← All Articles Who Noticed

Who Noticed

By Fernando Medrano

The argument about AI in family history research is conducted as a contest. Can the machine be trusted, yes or no. One camp posts about the remarkable thing an AI just did with their great-grandmother's letters; the other points out that it is confidently producing text which is not there. I have written elsewhere that the skeptics are right about transcription, and I still think so. But the contest framing has a defect that only shows up once you have done the work for a while: it forces every finding into a question about credit. Who found this? Was it the AI or was it you? And that question, applied to real examples, turns out to be either unanswerable or uninteresting. What follows is four things this project found, and an honest account of where each one came from. They are not arranged to show that the machine is clever, or that it isn't. They are arranged from most human to most automatic, because the interesting thing is how little that ordering explains. The useful distinction, when it finally arrived, was not between the machine's work and mine. It was between two entirely different questions that both get called judgement — one of which a machine can answer, and one of which it cannot.

A hunch, a count, and the letter that complicated both

The archive spans about forty years, and I had a suspicion that the language of the letters would track the family's assimilation. Not a sophisticated idea — an immigrant family writes home in Yiddish, and their children write in English. But it is the kind of thing that is either visible in the data or is not, and I had no idea which. The machine answered in about a second. Sorted by decade, the letters that carry a date and a language read like this: in the 1900s, twenty in Yiddish against four in English. In the 1910s, fourteen English and no Yiddish at all. In the 1920s, thirty-four English against four Yiddish. The last decade in the archive is English only. That is a near-total inversion inside twenty years, and it is the sort of finding the contest framing handles badly. I supplied a hunch that a person raised on immigrant-family stories would have. The machine supplied a count that would have taken an afternoon by hand and was exactly as reliable as the metadata behind it. Neither of us did the interesting part. The interesting part was in a single letter, and the machine found that one too — and, unlike the count, it knew immediately that it had found something. During a transcription pass, an agent working through one of the English letters flagged a passage it could not read, said what kind of thing it was — a Yiddish phrase, here, unreadable — and brought it to me straight away as interesting rather than filing it with the rest. That distinction matters. There are thousands of unreadable-span marks in this archive and almost all of them are dull; the honest-uncertainty machinery produces them by the yard. This one came with a reason attached. The letter is from one of Sam's daughters, written from a boarding house in Bournemouth at Christmas. It is completely fluent English, and it is wonderful — she has been playing pontoon with her sister and winning, her mother has bargained a shopkeeper down on a silver thimble, and she would like her father to remember that he promised the girls five pounds each. Mid-paragraph, addressing him directly, she drops in a Yiddish idiom. She writes it phonetically, in Latin letters, because she can say it and cannot spell it in Hebrew script. A person reading that sees the whole assimilation arc collapsed into one sentence — and running backwards from what the count implies. Not the immigrant father hiding his Yiddish. The English-speaking child reaching back for it, and not quite reaching. The count files that letter as English. Correctly, by its own rule: it is an English letter. So the measurement is right and thin at the same time, and the thinness is not visible from inside the measurement. Nothing in the count could have told me that one of those thirty-four English letters contains a girl speaking Yiddish she cannot write. The hunch was mine. The count was the machine's. The flag on the unreadable phrase was the machine's, and so was the judgement that it was worth interrupting me for. What that phrase meant for the arc was mine. I do not know how to divide that up, and I no longer think dividing it up is a useful exercise.

Five notes that were already an argument

Sam went from twenty-five cents a day to printed business letterhead in about fifteen years, and no single document says so. The machine held every piece of that for weeks; a person asked the question that assembled them. Sam arrives in Montreal in 1908 as a trained furrier and cannot work as one. He has the skill and none of the three things that actually let a furrier trade — capital, local reputation, a supplier network. So he takes casual gang labour at roughly twenty-five cents a day, a fraction of what was already considered a bare subsistence wage, and writes home that life there is very dry. By the 1920s he is writing to his wife on his own printed business letterhead from hotels in Leipzig and Paris, and the archive's later documents are cheques, invoices and tax notices at a scale that has nothing to do with gang labour. Every piece of that was already recorded. There was a note on Montreal wages in their period context. A note on why he took grunt work despite the training. A note on the fur-trade margins implied by one Paris letter about moleskins. A note on the engagement contract, which specifies a dowry. Four or five separate observations, each written at a different time, none of them making a claim about the arc. What was missing was somebody asking the question. Paul asked how his great-grandfather's fortunes changed, and the answer is a synthesis that no single note contained: the capital problem, the return to London where his network already existed, and — the genuinely new link — the dowry as plausible seed capital, which nobody had connected to the trading he was suddenly able to do. That is not the machine finding a pattern. It is a person asking a question of material that had been sitting there, in a form that made the answer assemblable. The work the machine did was earlier and less glamorous: it wrote things down, in a place where somebody would find them.

A question that needed an instrument built for something else

Some questions cannot be asked until an instrument exists, and cannot be answered once it does without a person to ask. This is the case I keep returning to, because I cannot tell a version of it where either party works alone. On the thirty-first of July, we built a confusion matrix: a model of which Hebrew letter shapes these handwriting models mistake for each other, derived from thousands of places where two independent readings of the same page disagree. It exists to answer a narrow question — where should we look for transcription errors. Its top row says that one particular pair, a tsadi read where the ink probably says a dalet, accounts for more confusions than any other. A week later I was looking at something unrelated: a person in the archive whose name is recorded, in the data, as a bracketed guess with a question mark in it. The kind of record that exists because the extraction was honest about not knowing. And I asked whether that name might be a different word entirely — whether the first character was one of the two the matrix says are most often confused. That question is not available to a person without the matrix. I had no independent reason to suspect that particular substitution; I knew about it because a machine had measured it a week earlier for another purpose. And the question is not answerable without a machine either — checking it meant comparing every independent reading of every Yiddish letter in the archive, which is not an afternoon's work by hand. The answer, when the sweep ran, was that three models read that greeting and two of them say tsadi while the third says dalet, differing on exactly the character the matrix ranks first. Reading it as a dalet makes it a diminutive of the recipient's own name — a woman who appears in about a hundred and fifty letters — rather than a stranger who appears once and never again. The same sweep found the same disagreement in three other letters, which reopened a question about a nickname that had been sitting unresolved for a month. The argument that had kept it unresolved was that one letter showed the name with no competing reading anywhere. That was true when it was written and had quietly stopped being true, because a fourth transcription model had been added in the meantime and nobody had gone back to check. None of this is a machine being clever, and none of it is me being clever. It is an instrument built for one purpose being available when a different question showed up.

The one where the numbers held and a person said no

A machine proposed the most affecting sentence anyone could write about this collection, and a person rejected it. That is the closest thing here to a machine having an idea, and it changed how I think about the whole arrangement. An extraction pass over the archive produced a proposed theme, stated as a fact about the collection: of the roughly hundred and thirty legible letters, not one is written by Dora. She is the addressee of most of the archive, the subject of constant concern, the recipient of the money and the reassurance and the anxious questions about her health — and the correspondence that survives is almost entirely other people's words about her, or to her, and never once from her. That is true. It is trivially checkable, and it was checked. It is also, I think, the most affecting single sentence anyone could write about this collection. It was rejected as a theme. The rejection is in the data, with a name and a date on it, sitting next to the twelve themes that were accepted. When I first saw that I assumed it was an oversight. It is not. The theme records in this project carry two separate rulings, and the code that defines them is unusually blunt about why. The first field answers exactly one question: are the quotes this thesis cites actually real? It explicitly does not say the thesis is established. The second is described in the source as the second, independent ruling — is this an editorial reading rather than a fact about the corpus? The split exists because one gate was not enough. Everything passed it. Theses where the man says the thing outright, in his own words, in a letter, came back verified. So did theses that are pure interpretation over an absence — including this one, which carries zero pieces of supporting evidence, because its entire claim is that something is missing. The same one-question gate approved both, and approving both is how a collection of hunches turns into something that looks like a corpus of established facts. So the numbers held perfectly and a person refused it anyway, on the grounds that a true and moving observation about an absence is an interpretation, and interpretations get labelled as such.

The record is biased toward the machine

There is a problem with everything above, and it is worth naming because it will affect anyone who tries to reconstruct how a project like this actually went. The machine leaves artifacts. Every agent pass leaves a memo. Every count leaves a query and a result. Every change leaves a commit with a timestamp on it. A person's hunch, spoken out loud in the middle of a conversation, leaves nothing at all. So if you reconstruct who noticed what from the written record — which is the only durable source — you will systematically over-credit the machine. Not because it did more, but because it is the only party that writes down what it was thinking as a matter of course. This is not hypothetical, and I ran into it assembling the examples above. Going back through the files to work out where the language-shift finding came from, the record shows a count and an agent-flagged phrase and nothing else — it reads as though the machine found the whole thing. The hunch that prompted the count left no trace at all, because it was a sentence somebody said out loud. I only have it right here because I was the one who said it, and I happened to remember. And then, editing this article, I found the same bias running the other way. I had written that the agent flagged that phrase without knowing what it had — the record shows a terse mark on an unreadable span, which is what thousands of dull ones look like. That was wrong. It brought the passage to me immediately, as something interesting, and I only know because I was there. The artifact makes a judgement look mechanical the same way the absence of an artifact makes a hunch disappear. Neither error is visible from the record, and both of them favour whichever party you already expected to be doing the work. Which brings me to the part of this project that does the most work and looks the least impressive. There are now something over two hundred research notes in this archive, written by both of us in about five weeks, over the summer of 2026. Most are not findings. They are hunches and to-dos: this might be a pattern, someone should check this, here is a piece of context that might matter later. They are cheap to write and deliberately assert nothing — a note records that somebody thought something was worth a look, and no more than that. Each one is tagged by what kind of thing it is: a topic, a task, a fix, a story, a person, a decision, a lead, an open question. Twelve of them are currently tagged as questions, which is a hunch that admits it is waiting for someone else. The important property is that the person who writes one is usually not the person who uses it. There is a note from the tenth of July about Yiddish diminutives — recording, for no particular reason, that a certain ending turns a name into an affectionate form of itself. Nobody wrote it in service of anything. A month later, when the question about that bracketed name came up, it was already there. That is the whole mechanism. A hunch written down stops belonging to the person who had it. It survives the session, and somebody who never had it can pick it up and act on it — including a machine, and including a person a month later who has forgotten having the thought.

Two questions that both get called judgement

Two separate things have to happen before a pattern can be treated as true, and only one of them can be automated. This is the distinction I wish I had started with. When a pattern shows up — from a count, from a hunch, from an agent flagging something it could not read — two entirely separate things have to happen before it can be treated as true. They get collapsed together under the word judgement, and collapsing them is what produces both the credulous version of this work and the dismissive one.
  • The first question is whether the numbers hold. Does the quote exist in the letter it is attributed to? Does the count reproduce? Is the artifact it cites actually in the archive? This is real work, it catches real errors, and a machine can do most of it — better than a person, in fact, because it does not get bored and does not stop at the point where the answer confirms what it expected.
  • The second question is whether it matters. Is this worth saying? Is it an observation about the collection, or a reading of it? Does asserting it as a fact overstate what the evidence carries?

Nothing checks the second one. There is no instrument, and there is not going to be. The theme about Dora's silence passes the first gate absolutely and fails the second, and both of those results are correct.

So the shape of the collaboration is not that the machine does the mechanical part and the human does the thinking. It is that noticing can come from anywhere — and did, in all four of these — while the two things we call judgement have completely different characters. One is a checkable property of the data. The other is a decision about what is worth claiming, made by someone who has to live with the claim.

AI or human was always the wrong question. The better one is which gate you are standing at, and whether you have noticed that there are two.