Artificial intelligence is already used in academic research for literature searching, title and abstract screening, full-text assessment, information extraction, and quantitative analysis. Large language models (LLMs) can process research text rapidly, but published evaluations show substantial differences between AI and human performance depending on the review task.
AI has a measurable advantage in processing speed. A 2024 study published in BMC Medical Research Methodology evaluated ChatGPT on 1,198 abstracts from three radiology fields. ChatGPT completed screening in less than one hour, while general physicians required an average of seven to ten days. The model achieved 95% sensitivity and a 99% negative predictive value.
The expansion of commercial AI services has also affected internet infrastructure and branding. The .ai extension is Anguilla’s country-code top-level domain, although registrations are available internationally. Organizations developing AI-related research tools or services can buy .ai domain registrations for AI-focused websites.
AI Can Process Literature Faster Than Human Reviewers
Literature screening is one of the most measurable applications of AI in evidence synthesis. Automated systems can evaluate titles and abstracts against predefined eligibility criteria without reading documents sequentially at human reading speeds.
Research has produced several quantitative findings:
- In the 2024 radiology screening study involving 1,198 abstracts, ChatGPT produced workload savings ranging from 40% to 83%.
- A 2025 diagnostic-accuracy study evaluated ChatGPT 4.0, Claude 3.5, Gemini 1.5, DeepSeek-V3, and RobotSearch on 1,000 publications.
- Mean screening time per article in that study was 1.3 seconds for ChatGPT, 6.0 seconds for Claude, 1.2 seconds for Gemini, and 2.6 seconds for DeepSeek.
- None of the five evaluated systems was considered suitable as a standalone literature-screening solution.
These measurements demonstrate that speed and independent reliability are separate performance characteristics. AI can reduce screening time without achieving sufficient accuracy to eliminate human verification.
AI Accuracy Varies Across Review Tasks
Systematic reviews contain several distinct tasks, including searching, screening, extracting data, evaluating study characteristics, synthesizing evidence, and calculating pooled statistical estimates.
A study comparing ChatGPT with human-conducted chronic-pain systematic reviews reported:
- 70.4% accuracy in title and abstract screening.
- 54.9% sensitivity in title and abstract screening.
- 80.1% specificity in title and abstract screening.
- 68.4% accuracy in full-text screening.
- 75.6% sensitivity in full-text screening.
- 66.8% specificity in full-text screening.
Performance was considerably stronger in specific quantitative tasks. ChatGPT successfully pooled data for five forest plots and achieved 100% accuracy in calculations of pooled mean differences, 95% confidence intervals, and most heterogeneity estimates, although minor discrepancies occurred in tau-squared values.
Reference Hallucination Remains a Documented Limitation
AI-generated references create a specific reliability problem for literature reviews. A 2024 comparative study examined 471 references generated by GPT-3.5, GPT-4, and Bard while attempting to reproduce human systematic-review searches.
The study recorded hallucination rates of:
- 39.6% for GPT-3.5.
- 28.6% for GPT-4.
- 91.4% for Bard.
Precision was 9.4% for GPT-3.5 and 13.4% for GPT-4, while Bard recorded 0% precision in that experiment. The researchers concluded that the evaluated LLMs should not serve as the primary or exclusive method for conducting systematic reviews.
Another 2024 evaluation found that only seven of 1,287 studies identified by ChatGPT were directly relevant to the review topic. The human benchmark contained 24 relevant studies.
Human Reviewers Retain Advantages in Complete Systematic Reviews
A 2026 Scientific Reports study compared six LLMs with human researchers across literature searching, screening, data extraction, analysis, and final systematic-review drafting.
The reference systematic review contained 18 articles. The best-performing LLM retrieved 13 of them. Data extraction and analysis were only partially accurate, and generated reviews frequently failed to comply completely with the PRISMA 2020 reporting structure.
A separate 2026 study tested ChatGPT 4.0 and 5.0 on 170 full-text articles concerning influenza-vaccine effectiveness. ChatGPT 4.0 achieved 71% accuracy, while ChatGPT 5.0 reached 77%. ChatGPT 5.0 recorded sensitivity of 87% and specificity of 70%. Its repeated decisions showed 80% agreement, demonstrating that identical review tasks did not always produce identical model decisions.
Human-AI Collaboration Has Measurable Advantages
Research outside systematic reviewing also provides evidence about human-AI collaboration. Workplace experiments summarized in research on creative human-AI teams include a study of 776 professionals at Procter & Gamble. Individuals using GPT-4 produced solutions of comparable quality to human teams without AI, while human teams using AI generated the highest-quality solutions. AI-assisted participants also completed assigned tasks approximately 12–16% faster.
The same principle is supported by literature-review research: AI can perform high-volume processing while humans verify eligibility decisions, references, extracted information, methodological judgments, and interpretations.
Can AI Actually Become Better Than Humans?
Current evidence does not establish AI as a superior independent literature reviewer across the complete review process. It establishes narrower areas in which AI already exceeds human capabilities, particularly processing speed and the rapid handling of large numbers of documents.
The documented evidence supports three distinctions:
- AI can screen research documents substantially faster than human reviewers.
- AI can perform some structured calculations with high accuracy.
- AI still produces false exclusions, incorrect selections, inconsistent decisions, and fabricated or inaccurate references.
Published comparisons therefore support AI-assisted literature review rather than fully autonomous literature review. Current systems can reduce human workload and accelerate individual stages, but human verification remains necessary for reliable systematic evidence synthesis.

More Stories
Behind on Taxes in Washington DC? A First-Steps Guide to the IRS and the OTR
Can entertainment sites be fully automated through AI?
Inside The Technology Systems Driving The Next Generation Of Sports Betting