← Back to start How This Was Made
The process behind the archive — a real experiment in using AI to do genealogical research honestly.

A Box of Letters

How This Was Made

This project is two things at once: the recovered story of Sam and Dora, and a real, honest test of whether AI tools can be trusted to help with family history research at all. No one left in the family reads Yiddish well enough to check this alone — so here's exactly how we've tried to keep anything from getting made up, including the parts of that effort that are still in progress.

AI
This site is a Human + AI Collaboration

This whole page is the answer to that question — read on for the full account of how AI and a human worked together to build this archive.

Phase One — Getting the Words Right
1
Scan it first
Every letter and postcard was scanned first, before anything else happened to it. That scan is the one fixed, unarguable thing in this whole project — every fact on this site has to trace back to it.
2
Read more than once, at every layer
The handwriting is now read by four independently-trained tools, none seeing another's answer first, and the resulting transcription is translated into English by two more. The point of the extra readings isn't more confidence — it's better-located doubt. Comparing them showed that our two newest tools agree with each other more closely than either agrees with the reading this site was originally built on, which tells us where that older reading most deserves a second look. It isn't proof the newer ones are right; machines agreeing with each other never is. It's a map of where to spend a person's time.
Why it matters: a lot of this correspondence is badly damaged — faded ink, torn pages, difficult cursive. Two tools reading the same handwriting independently and agreeing is real evidence a reading is right — that comparison has already caught a genuine, repeated transcription error in this project, one Hebrew letter shape mistaken for another across dozens of words. A fluent-sounding translation isn't the same kind of evidence: a translator can produce smooth English from a garbled transcription just as easily as from an accurate one.
3
Compare the answers — and say what we don't yet catch
Where the two Yiddish readings agree, that agreement is a real cross-check, and we trust it. Where they disagree, we show it — that disagreement is exactly how a real, systematic misreading was found. What this doesn't yet catch is a transcription that's confidently wrong but still looks like an ordinary, plausible word to every tool that reads it: it won't get flagged, because nothing about it looks unusual. Closing that gap is ongoing work, and it's why we're preparing to commission a professional Yiddish transcriber, focused on exactly the letters where the two machine readings disagree the most. That was written as a caution, and it earned its place: the method has since caught exactly that kind of reading twice — a garbled passage turned into fluent, confident English. Both are corrected, and both were found by the checks described here rather than by chance.
What this looks like on the page: phrases like “[illegible]” or a note that a name or date is uncertain are the honest record of what a tool couldn't read — not a guarantee that everything else on the page was read correctly.
4
A person checks it — and we learned how to spend that attention well
A family member worked through the earliest letters line by line, adding what a computer couldn't know: a family nickname, a local street, a relative's face in a photograph. Then something useful happened. He noticed that on a language he doesn't read, his own guesses about an unclear word weren't adding information — so rather than keep going letter by letter, he stopped, and left those passages standing as the machines had marked them. That judgement turned out to be exactly right, and it took a month to prove. The two readings we have since refined both came from those first heavily-edited letters, because that particular kind of slip needs a person confident enough to settle a question the evidence leaves open. It is the signature of careful work, not careless work. So here is where human attention actually sits: a small number of letters have had a full line-by-line pass; the rest carry the machines' own uncertainty markers, visible in the text, exactly where the reading is unsure. That is a deliberate division of labour rather than a gap — and it points a reader who can read Yiddish straight at the passages where their help is worth most. Two have already written in. It is the single most effective check this project has.
Phase Two — Turning Verified Letters Into a Story

Getting an honest translation is only half the job. The letters also had to become a story — connected across decades, explained where they assume things a modern reader wouldn't know, and honest about which parts are solid and which are educated guesses.

5
Put everything in order
Every letter and document was dated as precisely as the evidence allowed, then grouped into the eras you see on this site — the Courtship, the War Years, the Travelling Years — based on where Sam was actually writing from and what he was doing, not by assumption.
Where a date wasn't certain: the site says so directly — “c. 1907–1908” means a best estimate, not a known fact dressed up to look precise.
6
Bring in outside history, clearly labeled
Some of what makes the letters make sense isn't in the letters at all — what the Panic of 1907 actually was, how emigration routes worked, why a Yiddish-speaking family would leave the Pale of Settlement. That's real historical research, brought in on purpose and clearly marked as background — never blended into what the letters themselves say.
7
Look for patterns across the whole archive
Some of this site's biggest findings aren't from any single letter — they only show up when you count across all of them: how many surviving letters are in Sam's hand versus Dora's (none of hers survive, though Sam's own letters make clear she wrote back — "I received your letter" — so this is about what was kept, not what existed), how many open with some version of “no letter from you,” which years have almost nothing at all. These are measurable facts about the collection itself, not interpretation.
8
Flag the guesses as guesses
Where the story connects two things that no letter actually states outright — like whether the wedding dowry became the seed money for Sam's fur business — that connection is labeled as speculation, not fact. Family memory (a name from a family tree, a date recalled by a living relative) gets used too, but always held at a lower confidence than an actual document, and never silently upgraded into one.
The scan is ground truth
Every claim on this site — every date, every name, every quote — has to trace back to a scanned page, a family tree record, or named family testimony. If it can't, it doesn't get stated as fact.
Honesty over fluency
It would be easy to smooth over the gaps and write confident, tidy prose. We didn't. Where the letters are ambiguous, this site stays ambiguous too — uncertainty is preserved, not polished away. The machinery mostly holds: 82 letters are flagged uncertain outright, and thousands of markers in the text show where a word couldn't be read with confidence. Twice, a marked-uncertain passage was resolved into clean prose that a later reading refined — the honest limit of the principle, and the reason the markers stay visible rather than being tidied away.
Rigor builds the archive; a person tells the story
The full, honest record behind this site is exhaustive on purpose. But the story you actually read — what's featured, in what order, in whose voice — is a family member's own editorial retelling of that record, not a claim that anything left out is untrue.
PL
Paul Levine
Project originator · family archivist · design lead
A retired graphic designer who ran Tangent Design. Two earlier projects share a thread with this one — Barbed Wire, a beloved 1990s Vancouver webzine, and Raising the Dead (2002), a Canada Council–funded interactive memorial for his brother Richard — whose files, decades later, turned out to hold the box of letters this whole project is built from. On this project: supplied the letters and family testimony, reviews and approves every reading, holds final editorial authority, and personally designed this entire site.
FM
Fernando Medrano
Creative technologist
A Vancouver-based creative technologist, roughly twenty-five years into building the systems behind creative work — a decade at Radical Entertainment, then directing the art of Glitch for the studio that became Slack, then founding studios of his own. On this project: created the architecture and machinery that turns a box of fragile, half-legible pages into a structured, searchable archive; the rules that keep that archive honest; and the system that gives the material shape and meaning, then narrative — and ultimately an engaging, living account. One rule is built into its bones: it can never present a guess as a fact.
Under the hood — a more technical look, and what we learned

A longer, more technical account of how this archive works — the division of labour between the machines and the people, what the honesty checks actually check, and what we learned when they caught something. Alongside it, further writing about using AI in family history research: what we found, what broke along the way, and how the checks themselves are built.

  • Genealogy in the Age of AI — the argument: what a specialist told us, why the skeptics are right, and where the errors actually came from. Read
  • What We Measured, and What Broke — the findings: seven approaches that didn't work, and the instruments that lied to us about their own results. Read
  • How This Was Built — the technical account: the architecture, the tools, the day-to-day workflows, and enough to start your own. Read
  • The Machine That Refuses — how the honesty is enforced rather than merely intended: what the software declines to do, what the people promise, and why neither half works alone. Read
  • Who Noticed — on working with the machines rather than against them: four findings, where each actually came from, and the two different questions that both get called judgement. Read
  • Just Add Water — the transferable part, for anyone starting their own archive: the posture rather than the code, with a file to hand your own AI assistant. Read
  • A Beautiful Fabrication — the transcription models themselves: the same letters through four Yiddish handwriting models, what the published accuracy scores hide, the ones we rejected and why, and what the whole comparison cost. Read