← All Articles A Beautiful Fabrication

A Beautiful Fabrication

By Fernando Medrano

This archive holds about a hundred and fifty letters. Seventy-two of them are in Yiddish, most of those handwritten, and nobody in the family reads Yiddish. What follows is a field report on the AI models that read them instead: four handwriting models run over the same pages, the vision models we used alongside them, and the ones we rejected — with the dates, the numbers, and the evidence each verdict actually rests on.

Why there is no one left to read them

Before the war, Yiddish had around eleven million speakers. The YIVO Institute for Jewish Research counts 10,690,000, most of them in Eastern and Central Europe, with about three million more in North America. The Holocaust killed roughly half of them. Nothing else in this section is comparable to it. What happened afterwards was slower and had several causes at once. According to YIVO, Yiddish usage in the second half of the twentieth century was reduced by acculturation and assimilation in the United States, by forced acculturation and assimilation in the Soviet Union, and by repression of Yiddish and acculturation to Hebrew in the State of Israel. In Israel that meant an explicit national preference for Hebrew, publication limits on Yiddish newspapers enforced through licensing law inherited from the Mandate, and a permit regime restricting Yiddish theatre, concentrated around 1949 to 1951. It was not a single decree and not a total ban, and the historian who has written the fullest account of it argues the Hebrew-only project was only partly successful. It is one factor of three, operating on a population already halved. And for this family it was none of those things. They settled in London. What happened to their Yiddish is the most ordinary cause on the list: two generations, and the children spoke English. That is why the box came to us unreadable, and it is why the rest of this article is about software.

What we ran

Four handwriting models, all of them hosted on Transkribus, each identified there by a numeric id so you can look it up yourself.
  • The Dybbuk, model 46159. This was the incumbent — the model the whole site was originally built on. It was trained mainly on play manuscripts from a research project of the same name. The archive holds 74 transcriptions produced by it.1
  • Mame Loshn Maven, model 371445. Trained on 867 pages of civil records and letters combined — 33,487 lines, 139,504 words. It produced 75 transcriptions here.
  • Letter Reader, model 371705. Trained on 532 pages: family letters, relief-society correspondence, community minutes — 12,240 lines, 99,526 words. Also 75 transcriptions.
  • Yiddish to Latin, model 116033. This one does something different, and it is worth understanding because it explains the shape of the whole pipeline. It does not read a page into Yiddish. It reads the page into the Latin alphabet — the same Yiddish words, spelled phonetically in the letters this sentence is written in, so that lib̤entse dore is what the other models render as ליבענדע דארע. That makes it useful as an independent check rather than as a source: it reads the same ink through a different output alphabet, so where it disagrees with the others, the disagreement is about the strokes rather than about a shared habit of spelling. It produced 75 transcriptions.

Which is the moment to be plain about what the pipeline actually is.

Yiddish is written in the Hebrew alphabet. It is not Hebrew — it is a Germanic language with Hebrew, Slavic and Aramaic in it — but it uses those letters, and that is what these models produce: Yiddish words in Hebrew script. Nobody in this family reads either. So the published English on this site is not a translation of the scans. It is a translation of a transcription: a handwriting model reads the ink and produces Yiddish in Hebrew characters, and then a second model translates that text into English. Sixty-eight of the English translations in this archive cite the Dybbuk transcription as their source, not the photograph.

That is two steps, and the second one cannot see the page. Every misread character in step one arrives in step two as a fact, and a translator working from text has no way to know the ink said something else. It is why the character confusions further down this article matter more than they first appear: a single wrong letter does not stay a wrong letter. It becomes a wrong word, and then a fluent English sentence built on it.

Alongside the handwriting models, Claude's vision model produced 76 readings — nearly all of them on the English material, where the writing is printed or clear enough to photograph and read directly. That is the whole roster: 299 recorded jobs across the four handwriting models plus one that never finished, and 151 records in the archive, of which 72 are Yiddish, 76 English, and 3 mixed.

  1. Our own job log holds 73 entries for that batch, not 74. Both counts are correct: one early run predates the log entirely, so it left a transcription in the archive with no matching job record. Two different things counted honestly, rather than one thing counted twice.

What the scores don't tell you

A published accuracy score tells you almost nothing about how a model will read your pages. Each of these publishes one — character error rate, the share of characters it gets wrong — and lined up, they look like a ranking. The Dybbuk reports 4.4%. Mame Loshn Maven reports 7.68%. Letter Reader reports 9.35%. And there is a fourth Yiddish model out there, trained on synthetically generated handwriting, reporting an error rate of 0.0009 — which is to say, essentially perfect.
The same four figures, tabulated. Each is measured on that model's OWN held-out data, which is why they cannot be ranked against each other.
ModelPublished error rateMeasured on
The Dybbuk4.4%Play manuscripts
Mame Loshn Maven7.68%Civil records and letters
Letter Reader9.35%Family letters, community minutes
A fourth, synthetic-trained model0.0009Synthetically generated handwriting

The ranking is worthless. Each figure was measured on that model's own held-out data — the pages its makers set aside during training so they could score it on writing it had never seen. The Dybbuk's 4.4% means it reads play manuscripts well, because play manuscripts are what it was tested on. Letter Reader's 9.35% is a worse number measured on a harder and much closer thing: actual family letters. On our pages, the model with the best published score was the outlier of the three — the one the other two most often disagreed with.

The synthetic model makes the point without ambiguity. A near-perfect score on synthetic data is evidence about the data, not about the model. If a program generates the handwriting and the same program grades the reading of it, the number tells you the two halves agree. It tells you nothing about a fountain pen moving fast across a thin sheet in 1908.

So the published error rate is not a property of the model. It is a property of the model paired with one particular test set, and unless that test set looks like your pages, the number is a fact about someone else's archive.

What we rejected, and on what evidence

We rejected three models, and the cheapest one cost a single page to rule out. Two of the three rejections went to the same model. On 2 July 2026 we tested Google's Gemini as a transcription source — the calibration ran against gemini-3.1-pro-preview and gemini-2.5-pro. We gave it a control page — one of the cleanest letters in the archive, only 1.4% of it marked unclear by any other reading. It returned an entirely different letter. Fluent, plausible Yiddish, correctly formatted, signed by a person who does not appear anywhere on the page. We ran the same scan again and got a second different letter, with a second different invented signatory. It was rejected as a transcription source that day. On 31 July we gave it a fair retest, this time against gemini-3.1-pro-preview alone, because both objections to the first test were reasonable. The first objection was that we had asked it to do the wrong job — transcribe a whole page rather than check a single line. The second was that we had handed it unpreprocessed pages instead of clean line crops. We fixed both: individual line crops, cut from the page, with the model asked only to compare a crop against an existing transcription and score the agreement. The crops themselves are worth a paragraph, because they were free and most people do not know they exist. A handwriting model does not only return text — it returns the shape it read each line from, as a polygon of coordinates on the image. That geometry is a byproduct of the transcription you already paid for, so you can cut every line out of the page as its own picture and put a specific reading next to the specific ink it came from. We recovered 2,168 line crops across the archive for zero additional credits, because the service stores and versions the transcripts it has already produced: running recognition again costs money, asking for the result you already have does not. It took some calibration to get there — lining crops up with the right line of text, and working out which scans had been silently rotated during processing. Expect to do the same on any corpus: neither is a Yiddish problem, and neither is visible until you look at a crop and see the wrong handwriting under the right words. It fabricated five out of five. Every return was a stock Yiddish letter formula — the epistolary phrases that open thousands of letters — described with specific, confident detail about handwriting it could not see. There is one line in the archive with external ground truth, read by a human who knows Yiddish. Against it the handwriting model scored 0.13 and this checker 0.77, on a scale where zero is a perfect match and one is no relationship to the correct text at all. Then we turned the randomness setting to zero — the dial that decides how much a model is allowed to vary between runs, where zero should mean the same input gives the same output every time — and ran it again. One line came back with a different fabrication than before. So we stopped, eight API calls in, and wrote it down as a dead end. That is the whole decision. A checker whose score is not stable cannot have a threshold set on it: any cutoff we tuned on one pass would not hold on the next, which disqualifies it regardless of how accurate it might be on a good day. The stage we had built for it is still in the codebase, unused, with a note at the top saying what it does and why we do not run it. Gemini was accepted for translation, and still does that job on this archive. That is not a contradiction, and it is the most useful thing we learned. Translation hands it fixed text — the Yiddish is already there, in characters, on the input. Transcription hands it pixels of a hand it cannot resolve. When we asked it to translate, it preserved the unreadable markers we had put in the text and refused to guess at them. When we asked it to read ink, it invented. The variable is not the company or the model. It is whether the task contains a perceptual gap. A perceptual gap invites confabulation; a visible-text task does not. The third rejection was another handwriting model, Pinkas Brody, model 59324. Mean divergence from our existing reading was 0.84, with zero lines scoring under 0.50 — uniform garbling, no structure anywhere. That verdict rests on a single page, and the job for that page never even completed. We are not claiming more than that: one page, one unfinished job, and a decision not to spend more on it. A garble is safer than a fabrication. When a model returns nonsense, you can see it failing — the text does not parse, the words are not words, and no reader is fooled for a second. When a model returns confident, fluent, correctly-formatted prose that is not on the page, nothing about it looks wrong. A faithful garble is more truthful than a beautiful fabrication, because the garble tells you its own error rate and the fabrication conceals it. In an archive that no family member can check, that difference is the entire safety margin. These are findings about these pages — faint cursive, mixed scripts, 1900 to 1942 — on the dates named. They are not general verdicts about any model.

What agreement is worth

You cannot check one model with another model. That is the single most useful thing in this article, and it is the thing people most often get backwards — including us, at the start. The reasoning is seductive. Run a second model over the same page, see where the two agree, treat agreement as confirmation. It fails for a reason that is obvious once stated and invisible before: two models agreeing on an illegible character is the same failure twice, not corroboration. Where the ink is genuinely unreadable, similar models trained on similar handwriting will misread it in similar ways, and they will do it confidently. The method goes quiet exactly where you most need it to speak. Our own numbers show the trap. Across every Yiddish letter all three models had read, the two new models agreed with each other on 62% of about 11,655 comparable tokens, and each agreed with the incumbent on 49%. That looks like two newcomers corroborating one another against an old, worse reading — until you know that one of them was trained on the other's data plus a body of civil records. A shared training distribution is doing part of the agreeing. Their agreement is weak evidence about the ink and strong evidence about their pedigree. So the useful signal is inverted. Agreement tells you almost nothing; disagreement tells you where to look — it marks the places where two overlapping models were pulled apart by something actually on the page. We stopped treating a second model as a checker and started treating it as a finder of contested spots, which is a different tool for a different job. What the disagreements did surface was a set of specific character confusions. One substitution turned up 23 times, another 21, and across 25 letters the analysis pointed at 59 probable lost characters — not in rare words, but in some of the commonest words in the language, which is where a single wrong letter does the most damage to a sentence's meaning. The first version of that analysis invented 20% of its own evidence. Two different Hebrew characters transliterate to the same Latin letter, and the code compared the transliterations, so it manufactured confusions that never happened. The tell was structural rather than statistical: the result was perfectly one-directional — every instance ran the same way, with no counter-examples at all. Real handwriting confusion is messy and goes both ways. A perfectly clean result is a bug's signature, and we caught it by finding the cleanliness suspicious rather than impressive. I include our own instrument error here because it is the same failure mode as everything else in this article, committed by us. The measurement that did hold was a blind trial. On a set where 48% of tokens differed between models and not one line out of 37 was identical, both candidates independently reproduced a human expert's corrections — a Yiddish specialist who had never seen these models and did not know they existed. That included reading the recipient's name correctly, in a place where the incumbent had it wrong. Enormous disagreement in aggregate, and exact agreement with a human on the points a human had checked. Amount of disagreement told us nothing. What sort of disagreement told us everything.

Yiddish models, and what they cost

If you have Yiddish documents, this is the field as we found it — what we ran, what we did not, and what the whole comparison cost. Models we ran and measured. The Dybbuk, 46159, published error rate 4.4%, trained largely on play manuscripts — the incumbent, and on our pages the outlier of the three. Mame Loshn Maven, 371445, published 7.68%, trained on civil records and letters. Letter Reader, 371705, published 9.35%, trained on family letters and community correspondence — the closest training material to what we have. A Latin-script model for the English and mixed pages. Two rejections: Gemini as transcription source and again as line-crop checker, and Pinkas Brody, 59324, on one page. Models we did not run. Each of these carries the same label, and I want it stated on every row rather than once at the top, because a list of names reads as a list of recommendations otherwise. Transcription Aid Model 1, 51946 — a sibling of the transliteration model, plausibly relevant. We did not test this. DiJeSt 3.0, 357765 — trained on printed text, which is very likely the wrong domain for handwriting. We did not test this. Vaybertaytsh — trained on eighteenth and nineteenth century print, not cursive. We did not test this. VILNISH — from a different lab, on different data, which makes it the most genuinely independent candidate available and the one whose agreement with our models would actually mean something. We did not test this either. It is the most interesting name on the list and we have no measurement of it at all. And the synthetic-trained model with the 0.0009 score, which is a fact about its training data. We did not test this. Do this in order, because the first step is free and the order is what keeps it cheap. Start free. Transkribus gives 50 free credits a month with no card required. At one credit per page per model, that is enough to put a handful of your own pages through every candidate before you pay anything. Do this first. The whole point of the free tier is that it lets you disqualify models on your own handwriting, which is the only test that counts. Then about forty euros, total. Two months of the Scholar subscription at €19.99 each, which is €39.98. Scholar gives 150 credits a month, which at one credit per page per model is 75 letters through two models. That was the entire model-comparison budget for this project. The bulk runs went through the API rather than the web interface. API credits cost roughly half as much, and more importantly it is the only practical way to run hundreds of jobs and keep a record of each one.

Your own box of letters

Yiddish is the instance. The class is an archive in a language its own family can no longer read, and it is much bigger than one language. Ladino, carried out of Spain in 1492 and largely destroyed in the Balkans in the 1940s. Judeo-Arabic, ended by twentieth-century displacement across North Africa and the Middle East. Syriac, spoken by communities repeatedly scattered. Low German, which ordinary schooling and mobility replaced across northern Germany over a couple of generations. Rusyn, minoritised by one state after another. The causes are not the same — genocide, state policy, displacement, and the plain fact of children preferring the language school is taught in — but the outcome in a family's hands is identical. A box of paper, and nobody at the table who can read it. The method transfers even where nothing else does. Run your own pages through every candidate rather than trusting a published score, because the score was measured on somebody else's handwriting. Use the free tier to disqualify before you buy anything. Use a second model to find contested spots, not to confirm the first one — where they agree on unreadable ink, they are usually wrong together. Distrust any model that produces fluent text where the page is damaged — a torn or faded passage should come back broken, and if it comes back beautiful, something has been invented. And keep the scan permanently beside the transcription, so that anyone who does come along later, and can read it, can check. The models you need may already exist, built by people like you. Transkribus's public model list runs to 451 models. Only a handful of those are Yiddish. The two best models on our pages were released in 2025 by a community consortium of genealogists and archivists — not by a technology company, and not by anyone who was ever going to market them to us. And we found out they existed because a scholar mentioned them to a family member on Facebook. That is the entire discovery mechanism. There was no search that would have turned them up. I can say nothing about what exists for any language we did not check, and I am not going to guess. For a great many languages there will be nothing usable yet — no model, no training data, nobody working on it. That is a real and common outcome, and the free tier is how you find it out cheaply, in an afternoon, instead of after a purchase. So: search the Transkribus model list for your language, and ask the people who care about it. Those are two different actions and the second one found ours.