Ancient DNA is the most exciting thing to happen to genealogy in decades, and the easiest to misread. A single point on a chart, labelled with a place and a date, looks like a fact. Sometimes it is. Sometimes it is a sample that barely passed quality control, placed on the chart by software guessing at data it didn't have.
Here is a worked example, and a checklist.
A worked example: four ancient samples on one branch
A commercial Y-DNA site listed four ancient samples inside one paternal branch. For anyone carrying that branch, they looked like ancestors — men from a Roman-era study of the Levant, sitting right on their line.
Then we checked.
- None of the four is in the curated scientific release. The standard curated dataset contains 638 individuals from that study. These four are not among them.
- They survive only in a community-compiled spreadsheet of coordinates, and one is flagged there as pre-QC — recorded before quality control.
- Three of the four plot closest to Iron Age Poland and early medieval Croatia, not the Levant. The fourth sits near Imperial Roman Sicily and Iron Age Lebanon.
- Their distance from the only Jewish reference groups available was about as large as the distance between two random samples in the whole dataset.
- Their coordinates are imputed — filled in statistically — which drags low-coverage ancient samples towards the structure of the reference panel.
So either they really were Central Europeans, or their data is too thin to say anything. Both readings rule them out as evidence of Levantine descent.
A second example: the labels
You may have seen "Galilean Roman Jewish samples" circulating online. In the
actual published data, those individuals are labelled Israel_MLBA and
Israel_LateC — Middle-to-Late Bronze Age and Late Chalcolithic. That is
two to four thousand years before the Roman period. The label on the internet
was not the authors' label.
The checklist
Every time an ancient sample is offered as evidence, check:
- The assessment. Curated releases grade each sample — PASS, Questionable, Questionable_Critical. Know which one you're looking at.
- Coverage and SNP count. A sample read at a few tens of thousands of positions cannot carry a fine-grained conclusion.
- Curated or compiled? Is it in the official release, or only in a community sheet?
- Imputed or observed? Imputation is useful, and it also pulls weak samples towards what the software expects.
- Whose label? Is the group name the authors', or one someone gave it online?
And check the tree properly, too
The same care applies to reading modern trees. Some tree sites show ethnic and language labels only as hover text. Read the page quickly and you'll miss them. We made exactly that mistake and reported a branch as having no Jewish testers when it had a Libyan Jewish one. The fix is dull and reliable: look at the actual data, not the first impression of the page.
Takeaway
Anyone plotting a Questionable, low-coverage, pre-QC sample as a confident point on a chart is drawing a conclusion the data cannot carry. Treat ancient DNA like any other document: ask where it came from, how good the copy is, and who wrote the label.
Sources
- Allen Ancient DNA Resource (AADR) — curated annotation, ASSESSMENT column
- Akbari et al. (2026) data release and public annotation
- Community G25 coordinate compilations (cited as the source of the uncurated coordinates, not as evidence)