Research & PolicyFree
Remove Duplicates From a Zotero Library
Zotero calls two items duplicates on an identical DOI alone, so it offers unrelated articles and misses true duplicates in silence.
River reads the exports rather than asking you to tidy them first. RIS, BibTeX and EndNote XML all go in, from as many databases as you searched. Every record is then paired against every other on DOI, title, creators, year and item type. What comes back is a sheet of candidate pairs, each carrying its score, the fields that agreed, the fields that did not, and a recommendation you can argue with. Nothing is merged for you. The plan is the deliverable, so an interrupted run costs you nothing.
Everything ranking for this teaches the same four clicks. Zotero's own documentation is the best of them, and it says that when title, DOI or ISBN match, Zotero also compares years and creator lists. Its source at release 10.0.1 runs three independent passes, and the DOI pass calls the matcher with no comparison function at all. An identical DOI string is the whole test, with no title, year, author or item-type check.
Built for the reviewer who searched four databases and has a duplicates-removed count to report, and for anyone whose library has been imported into twice. Reach for it when the exports are in and screening has not started. What survives becomes the literature review paragraphs, then the thesis abstract. Interview transcripts are a different pile with the same counting problem, and that is research synthesis. Two saved reports disagreeing on one number is report reconciliation.
The count can be right while the set is wrong
Three unrelated papers in one journal can each carry the journal's own DOI rather than their own. Zotero's DOI pass selects every record whose DOI begins with the digits 10, folds the string to upper case, and unions everything that matches. Three records, one string, one duplicate set of three, and the merge button is live. That set exists because a field matched, not because the papers are the same work, and nothing on screen tells you which of the two it was.
Twenty-four records from four databases, eighteen of them distinct. The built-in detector offers six sets covering thirteen records, and refuses to merge one of them because the item types differ. Accept everything offered and it removes exactly six records, which is exactly the right number. Three of those six were distinct studies, and three true duplicates are still in the library. Twenty-four minus six is eighteen, the true unique count, reached by destroying three studies and missing three duplicates.
McKeown and Mir measured Zotero against a hand-checked benchmark of 3,130 references and found it 80 percent accurate, with 599 missed duplicates against 20 false ones. Missing a duplicate was close to thirty times more common than inventing one. Merging itself is safer than the forums suggest. The non-master goes to the Trash rather than being erased, notes and tags move to the survivor, and the whole action is undoable. The risk is the record, not the highlights.
How it works
Hand over the exports
Every export from every database, in whatever format your search interface gave you.
Pair and score
Each record against every other on DOI, title, creators, year, container and item type.
Write the reasons
Every pair gets a recommendation and the specific fields that produced it, including the rejections.
Work the queue
You decide the ambiguous pairs, then the count and the merge log come from your decisions.
What you get
- One row per candidate pair, with the fields that agreed and the ones that did not
- Pairs scored across item types, so a report and its journal article are compared
- A stated reason for each recommendation, including every pair the scorer rejected outright
- Field-level survivorship, naming which record's volume, pages and abstract should survive the merge
- A separate document for the ambiguous pairs, each with the fact that makes it ambiguous
- A duplicates-removed count that traces to the pairs behind it, not a library total
Common questions
Why is Zotero showing 53 items as duplicates when only two match?
Because they share a DOI. The DOI pass selects every record whose DOI begins with 10, uppercases it, and groups everything identical, with no title, year, author or item-type check. Chapters of one book and articles in one journal often carry the container's DOI, which is enough to bind them all into a single set.
How do I know how many duplicates were actually removed?
Count the pairs you accepted, not the items left in your library. A library total tells you what remains; it cannot tell you whether the right things went. The merge log carries one line per accepted pair, so the duplicates-removed figure you publish points at the decisions behind it. It is the discipline a results section needs later, where each figure names the line of output it came from.
Will I lose my annotations and notes if I merge?
No, and this is the fear worth putting down. Zotero moves the other copy's attachment, notes, tags and collection memberships onto the survivor, trashes rather than erases the loser, and stages the whole thing as one undoable action. When both PDFs carry embedded annotations it keeps both files rather than choosing.
Why will Zotero not let me merge two items it just called duplicates?
Because merged items must share an item type. An institutional report and the journal article that followed it will match on title and authors, get offered as a set, and then refuse to merge. Retyping one to clear the block is the wrong fix: under PRISMA those are two records, not one.
Which duplicates is Zotero missing?
The ones no screen shows you. An abridged title fails because the title test is exact equality after normalization. A matching ISBN on two book chapters is never queried, because that pass reads books only. And two copies carrying different DOIs are vetoed before their titles are ever compared.
Can I find duplicates in just one collection?
No. Zotero's detection works only within a library, never within a collection. Worth knowing that adding an item to several collections does not duplicate it, since one item can sit in many collections at once. Where a library really is several projects, filter the scored pair sheet rather than splitting the library.
Does this replace Zotero, or Covidence?
Neither. Zotero stays your library and Covidence stays your screening tool; what is missing between them is the pair sheet that says which merges are safe. Export RIS, work the queue, then merge in Zotero with the plan beside you. More on how researchers use River.
Remove Duplicates From a Zotero Library
Fill in the form and your workspace opens with the work already underway.