Abstract
Systematic reviews provide high-quality evidence but require extensive manual screening, making them time-consuming and costly. Recent advancements in general-purpose large language models (LLMs) have shown potential for automating this process. Unlike traditional machine learning, LLMs can classify studies based on natural language instructions without task-specific training data. This systematic review examines existing approaches that apply LLMs to automate the screening phase. Models used, prompting strategies, and evaluation datasets are analyzed, and the reported performance is compared in terms of sensitivity and workload reduction. While several approaches achieve sensitivity above 95%, none consistently reach the 99% threshold required for replacing human screening. The most effective models use ensemble strategies, calibration techniques, or advanced prompting rather than relying solely on the latest LLMs. However, generalizability remains uncertain due to dataset limitatio ns and the absence of standardized benchmarking. Key challenges in optimizing sensitivity are discussed, and the need for a comprehensive benchmark to enable direct comparison is emphasized. This review provides an overview of LLM-based screening automation, identifying gaps and outlining future directions for improving reliability and applicability in evidence synthesis.
| Original language | English |
|---|---|
| Title of host publication | Proceedings of the 14th International Conference on Data Science, Technology and Applications (DATA 2025) |
| Subtitle of host publication | Volume 1: DATA |
| Publisher | SciTePress |
| Pages | 508-517 |
| Number of pages | 10 |
| ISBN (Electronic) | 978-989-758-758-0 |
| DOIs | |
| Publication status | Published - 2025 |
| Event | 14th International Conference on Data Science, Technology and Applications, DATA 2025 - Bilbao, Spain Duration: 10 Jun 2025 → 12 Jun 2025 |
Conference
| Conference | 14th International Conference on Data Science, Technology and Applications, DATA 2025 |
|---|---|
| Country/Territory | Spain |
| City | Bilbao |
| Period | 10/06/25 → 12/06/25 |
Keywords
- systematic review
Fields of Expertise
- Information, Communication & Computing
Fingerprint
Dive into the research topics of 'Evaluating Large Language Models for Literature Screening: A Systematic Review of Sensitivity and Workload Reduction'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS