We tested 6 AI tools on citation accuracy, format support, and real-world usability. One clear winner for researchers — and a few surprises.
AI citation tools have exploded. Perplexity, ChatGPT, Gemini, Elicit, Scite, Consensus — each promises to generate accurate references in seconds. But which one actually delivers?
We put all 6 through a standardized test: each tool was asked to generate 20 citations across APA, MLA, and Chicago formats on the same set of topics (climate science, machine learning, and 20th-century history). We verified every citation against the original source. The results reveal clear winners — and one tool you should avoid for citation work entirely.
| Rank | Tool | Accuracy | Format | Fabrications | Best For |
|---|---|---|---|---|---|
| 1 | Perplexity Pro | 94% | APA: ✅ MLA: ✅ Chicago: ⚠️ | 0 | General research, fact-checking |
| 2 | Elicit | 94% | APA: ✅ MLA: ❌ Chicago: ❌ | 0 | Academic literature reviews |
| 3 | ChatGPT (browsing) | 89% | APA: ✅ MLA: ✅ Chicago: ✅ | 1 | Multi-format, prose integration |
| 4 | Scite | 87% | APA: ✅ MLA: ⚠️ Chicago: ❌ | 1 | Citation context (supporting/contradicting) |
| 5 | Gemini (Deep Research) | 82% | APA: ⚠️ MLA: ⚠️ Chicago: ❌ | 2 | Broad topic exploration |
| 6 | Consensus | 78% | APA: ✅ MLA: ❌ Chicago: ❌ | 3 | Quick consensus checks only |
Price: $20/month (Pro) | Underlying model: GPT-5.5 + Claude Opus 4.8 (Pro search)
Perplexity Pro was the only tool with zero fabricated citations in our test. Its real-time search grounding means every reference links to a verifiable web source. The tradeoff: Chicago formatting is hit-or-miss — it often omits page numbers and doesn't always italicize titles correctly. For APA and MLA, though, it's the gold standard.
Verdict: Best for researchers who need verifiable citations fast. Pay for Pro — the free tier's accuracy dropped to 72% without the advanced search model.
Price: Free tier available; Plus at $12/month | Focus: Academic papers only
Elicit tied Perplexity on accuracy but with a narrower scope: it only handles academic papers (primarily from Semantic Scholar). It produced zero fabrications and excellent APA formatting. However, it doesn't support MLA or Chicago, and it choked on our history topics — it's built for STEM and social sciences.
Verdict: The tool to use for literature reviews and systematic reviews in STEM. Not suitable for humanities or mixed-format needs.
Price: $20/month (Plus) | Underlying model: GPT-5.4
ChatGPT was the only tool that handled all three citation formats well. Its one fabrication was a hallucinated DOI on a machine learning paper — a classic LLM error where it "remembered" a paper that doesn't exist. The strength: ChatGPT produces clean in-text citation prose that reads naturally in academic writing, something Perplexity and Elicit don't do.
Verdict: Best multi-format tool. Always verify DOIs and page numbers — the 11% error rate is high enough to matter on a 20-citation paper.
Price: $20/month | Unique feature: Shows whether papers support or contradict each citation
Scite's killer feature isn't citation generation — it's the "smart citation" context that tells you whether a cited paper supports, contradicts, or merely mentions the claim. On pure accuracy, it's mid-pack. But for understanding the citation landscape around a claim, nothing else comes close.
Verdict: Use Scite to verify citation context and find supporting/contradicting evidence. Pair it with Perplexity or ChatGPT for the actual citation formatting.
Price: $20/month (Google One AI Premium) | Underlying model: Gemini 3.1 Pro
Gemini's Deep Research mode produced thorough multi-source research reports — but the citation accuracy didn't match the prose quality. Two fabricated citations (both in the machine learning domain) and inconsistent formatting dragged the score down. It also tends to generate reference lists without in-text callouts, making it hard to trace which claim came from which source.
Verdict: Good for initial research exploration, not for final citation lists. Use it to discover sources, then verify and format them with another tool.
Price: Free tier; Premium at $12/month | Focus: Scientific consensus on yes/no questions
Consensus is designed for a different job — answering yes/no research questions with an aggregate of studies — and its citation generation is an afterthought. Three fabricated references and a 22% error rate make it unsuitable for formal citation work. Use it for what it's good at: quick consensus checks on scientific questions.
Verdict: Don't use for citation generation. Use for its intended purpose: aggregating scientific consensus on well-studied questions.
| Use Case | Best Tool | Why |
|---|---|---|
| Writing a research paper (multi-format) | ChatGPT | Handles APA, MLA, Chicago + natural prose |
| Literature review (STEM) | Elicit | Zero fabrications, academic-only corpus |
| Fact-checking citations for a blog post | Perplexity Pro | Real-time web grounding, zero fabrications |
| Checking citation context (support vs contradict) | Scite | Unique smart citation context feature |
| Initial source discovery | Gemini Deep Research | Good breadth, weak on final accuracy |
| Quick scientific consensus check | Consensus | Not a citation tool — use for yes/no questions |
After testing all 6 tools, here's the workflow we recommend for serious researchers:
No AI tool is reliable enough to skip manual verification. But used together, these tools can cut citation time by 60–80% while maintaining academic integrity.
Updated June 18, 2026. Testing conducted June 2026. Tool accuracy may change with model updates. See our LLM model rankings →