Can a small model learn
from reading new papers?
MedAdapt Lab continues the training of a small open language model on research papers about ADHD that were published after the model was made, then tests whether it can answer questions about those papers — and what it forgets in exchange. Everything runs on one Apple Silicon Mac.
The question
Language models are trained once on a snapshot of text and then frozen. When a field publishes new findings, can a model pick them up just by reading the papers — and how much of its general ability does it lose on the way?
Does it remember specific findings?
Not “sound more medical”, but recall what a particular study found, compared with studies it never read.
Does the field become familiar?
Does new, unseen writing in the same field become easier for the model to predict?
What does it cost?
How much worse does the model get at ordinary, non-medical text after the training?
Design
The key idea is a built-in control: half of the new papers are never used for training, and the exam asks about both halves.
Collect papers the model cannot know
Openly licensed (CC0, CC BY, CC BY-SA) ADHD papers first published in 2026, after the base model was released, plus older papers for general domain text.
Split them once, at random
Half of the 2026 papers may be used for training (group A). The other half is never trained on (group B). The split is fixed before anything else happens.
Write and freeze the exams
Multiple-choice questions from both halves, each tied to a sentence in its paper, written without knowing which half a paper is in. Exams are frozen by checksum.
Train, then compare
Take the exams before and after each training run. Learning a specific paper shows up as group A improving more than group B.
What is measured
Every run sits the same frozen exams before and after training; every item is kept so changes can be traced question by question.
| Measure | What it tells us |
|---|---|
| Group A questions | Recall of findings from papers the model was trained on |
| Group B questions | The control: same kind of questions about papers it never read |
| Unseen 2026 paper text | Familiarity with the field (perplexity on held-out papers) |
| General text | Forgetting (perplexity on ordinary encyclopedic text) |
| Existing medical question sets | Whether anything transfers to broader medical questions |
Changes are compared item by item, with confidence intervals and a minimum number of questions before any change is called significant.
Setup
Model and training
Small open base models (Qwen3, 0.6B and 1.7B parameters) adapted with LoRA, using a hand-written training loop on Apple's MLX framework. Model revisions are pinned.
Data
Full texts from the PubMed Central open-access subset, accepted paper by paper only when the licence allows it. Papers and derived data are not redistributed.
Status
Where the experiment series stands. This list is updated as work progresses.
- Data and exams — 539 licensed papers from 2026 split once into trained (269) and never-trained (270) halves; 1,478 questions written and frozen.
- First training series — baseline exams and runs on the 0.6B model, including a measurement fault found and fixed (below).
- Follow-up experiments — a second seed, rewritten papers, the 1.7B model and the amount of reading.
- Fictional trials — a cleaner test with 400 invented trials the model could not know or guess.
- Final write-up — the results below, including what did not work.
Results
In one paragraph
A small model does pick up specific findings from papers it reads, but only a little at a time. It needs many exposures, and a larger model holds more. Reading the same fact in different words works better than reading one text again. Every bit of learning cost some general ability. When the new material was narrow and repetitive, the cost was severe.
Five experiments
ADHD-01: first runs and a measurement fault. ADHD-02: rewritten papers. ADHD-03: a larger model. ADHD-04: how much reading. ADHD-05: fictional trials, a clean test of facts the model cannot know.
Findings 1–5 come from the real 2026 papers. LoRA adaptation of Qwen3-0.6B-Base and Qwen3-1.7B-Base on the 269 trained papers. “Extra gain” is how much more accuracy rose on questions about trained papers than on questions about never-trained papers. It is the part of the improvement that can be credited to reading specific papers. Brackets are 95% intervals from resampling whole papers. Most runs use a single training seed, so treat these as strong trends, not settled facts.
1. Reading the papers leaves a small, specific trace
After reading each trained paper about 1.65 times, both groups of questions improved: the model got generally better at this kind of paper. Questions about the papers it actually read improved more:
| Run | Trained-paper gain | Never-trained gain | Extra gain |
|---|---|---|---|
| 0.6B, seed 42 | +8.6 points | +4.5 points | +4.2 [+0.5, +8.0] |
| 0.6B, seed 43 | +9.5 points | +5.7 points | +3.8 [−0.2, +7.7] |
| 1.7B, seed 42 | +7.7 points | +2.6 points | +5.0 [+1.1, +8.8] |
Before training, the models already answered 33–38% of these four-option questions correctly, well above the 25% of pure guessing. Some findings are guessable from general knowledge. This is why the next experiment uses invented trials.
2. More reading helps the larger model, not the smaller one
One long run per model, examined at several points. Exposure is the approximate number of passes over the trained papers.
| Exposure | 0.5 | 0.8 | 1.0 | 1.65 | 3.3 | 6.6 |
|---|---|---|---|---|---|---|
| Extra gain, 0.6B | +3.2 | +4.0 | +3.9 | +4.2 | ||
| Extra gain, 1.7B | +1.0 | +3.9 | +6.8 | +8.7 |
The 0.6B model levels off at about four points however long it reads. The 1.7B model keeps adding paper-specific answers, to +8.7 [+4.4, +12.9] points after about 3.3 passes. At 1.65 passes the two sizes looked alike. The difference only appears with more reading.
3. Specific recall and general benefit peak at different times
For the 1.7B model, the improvement on never-trained papers, text prediction on unseen 2026 ADHD papers and training validation loss were all best at about one pass, then got worse. Recall of the trained papers kept rising. Stopping training when validation loss is lowest, a common rule, would have ended it before most of that recall.
4. Every gain came with forgetting
| Exposure | 0.5 | 0.8 | 1.0 | 1.65 | 3.3 | 6.6 |
|---|---|---|---|---|---|---|
| General text perplexity, 0.6B | +10% | +19% | +53% | +123% | ||
| General text perplexity, 1.7B | +7% | +7% | +14% | +27% |
Perplexity on ordinary encyclopedic text rose at every point (higher is worse). With heavy reading the 0.6B model also became worse than before on unseen ADHD papers. It had overfitted the 269 papers it read.
5. A little rewriting did not help
Adding four short rewrites of each paper (about 14% of the training text) gave no extra gain over the original text alone: −0.4 [−3.2, +2.5] and −1.0 [−3.9, +2.0] points with two seeds. The rewrites covered most of the tested facts, so the likely limit is the amount of rewriting, not its content. The fictional-trial experiment tests rewording properly.
6. Facts it could not know: fictional trials
Real findings can partly be guessed, so before training the models already scored 33–38% on four options. For a clean test we generated 400 fictional drug trials with randomly drawn facts: country, population, dose schedule, result, side effect and so on. Each trial was taught in one of two ways at the same exposure. P: four differently written texts, each read once. R: one text read four times. C: trials never described anywhere, as a control. All three started at chance (25%).
| Run | P (four texts) | R (one text ×4) | C (never seen) | P minus R, gain |
|---|---|---|---|---|
| 0.6B, seed 42 | 43.7% | 39.8% | 23.9% | +4.5 [+0.4, +8.6] |
| 0.6B, seed 43 | 49.1% | 42.1% | 22.6% | +7.6 [+3.3, +11.9] |
| 1.7B, seed 42 | 51.3% | 43.3% | 24.0% | +6.3 [+1.9, +10.7] |
- The models did learn facts they could not have known. The control stayed at chance throughout.
- Learning was late and sudden. Almost nothing after one or two passes. Clear recall appeared between four and eight passes, by which point each fact had been seen 16–32 times.
- Different wording beats repetition in all three runs. This settles what finding 5 left open: the rewrites of real papers were too few, not useless.
- The cost was severe. Training on this small, repetitive corpus made general text 5–11 times less predictable (perplexity) after a single pass and 37–94 times by the end. The model drifted towards writing trial reports. Real papers had cost about 10%. Learning narrow new facts without mixing in general text badly damages a model.
The fictional texts were heavily templated (most phrasing recurs across trials), so this measures recall of templated statements, not of natural prose. The fictional data is not published.
What this adds up to
- Continued training can add specific knowledge to a small model, but slowly. Most early gains are general familiarity with the field, not specific facts.
- How much sticks depends on model size and on varied exposure. Repetition helps less than rewording.
- Signals that look healthy, such as validation loss, peak long before specific recall does.
- Forgetting is the constant price. It grows with exposure and grows faster when the new text is narrow, which is why practical systems mix in general data or retrieve facts instead of training them in.
What went wrong, and what we changed
- A broken attention sink. After the first run, perplexity appeared to have risen by 238–413%. The cause was not general damage. The model keeps a very large activation on the first token of its input and uses it as an “attention sink”. Training had disabled it whenever that first token was the end-of-text marker, which our perplexity test always put there. Starting every training window with that marker fixed it. The measurement now runs both with and without the marker, and the two agree.
- A biased score. One medical benchmark (yes / no / maybe answers) was scored in a way that favours longer answers. The larger model therefore almost never chose “no”, and a reported decline was an artifact. The smaller model answered “yes” to almost every question, so that benchmark told us nothing about it. Scores are now reported without the length adjustment for such answers.
- Data checks that caught real problems. Abstract-only records accepted as full papers, papers dated 2026 that were first published earlier, and benchmark questions that only mentioned ADHD in a wrong answer were found and removed before any training.
- Interrupted runs. The Mac's graphics watchdog stopped several training runs. Training now saves a checkpoint every 50 steps and can resume.
Limits
- Multiple-choice questions measure recall of findings, not clinical understanding, judgement or safety.
- Publication in 2026 lowers the chance that the base model saw a paper, but cannot rule out earlier preprints or reports of the same work.
- These are small models trained on one Mac. Results may differ for larger models or different training methods.
- The questions were written by an AI model from each paper and checked against the source text; they have not been validated by clinicians.
- Our full-parameter training control (updating all weights instead of LoRA) never completed reliably on this Mac, so no method comparison is made.
- Most comparisons rest on one training seed, and the same frozen exams were reused across experiments. The intervals describe question and paper sampling, not training variability.
Code and licence
The project's own code and documentation are open source under the Apache License 2.0. Papers, datasets, model weights and trained adapters are not part of the repository; they keep their own licences. See the licence page.