Skip to main navigation Skip to search Skip to main content

Evaluating Large Language Models for Literature Screening: A Systematic Review of Sensitivity and Workload Reduction

  • Elias Sandner
  • , Luca Fontana
  • , Kavita Kothari
  • , Andre Henriques
  • , Igor Jakovljevic
  • , Alice Simniceanu
  • , Andreas Wagner
  • , Christian Gütl

Research output: Chapter in Book/Report/Conference proceedingConference paperpeer-review

Abstract

Systematic reviews provide high-quality evidence but require extensive manual screening, making them time-consuming and costly. Recent advancements in general-purpose large language models (LLMs) have shown potential for automating this process. Unlike traditional machine learning, LLMs can classify studies based on natural language instructions without task-specific training data. This systematic review examines existing approaches that apply LLMs to automate the screening phase. Models used, prompting strategies, and evaluation datasets are analyzed, and the reported performance is compared in terms of sensitivity and workload reduction. While several approaches achieve sensitivity above 95%, none consistently reach the 99% threshold required for replacing human screening. The most effective models use ensemble strategies, calibration techniques, or advanced prompting rather than relying solely on the latest LLMs. However, generalizability remains uncertain due to dataset limitatio ns and the absence of standardized benchmarking. Key challenges in optimizing sensitivity are discussed, and the need for a comprehensive benchmark to enable direct comparison is emphasized. This review provides an overview of LLM-based screening automation, identifying gaps and outlining future directions for improving reliability and applicability in evidence synthesis.
Original languageEnglish
Title of host publicationProceedings of the 14th International Conference on Data Science, Technology and Applications (DATA 2025)
Subtitle of host publicationVolume 1: DATA
PublisherSciTePress
Pages508-517
Number of pages10
ISBN (Electronic)978-989-758-758-0
DOIs
Publication statusPublished - 2025
Event14th International Conference on Data Science, Technology and Applications, DATA 2025 - Bilbao, Spain
Duration: 10 Jun 202512 Jun 2025

Conference

Conference14th International Conference on Data Science, Technology and Applications, DATA 2025
Country/TerritorySpain
CityBilbao
Period10/06/2512/06/25

Keywords

  • systematic review

Fields of Expertise

  • Information, Communication & Computing

Fingerprint

Dive into the research topics of 'Evaluating Large Language Models for Literature Screening: A Systematic Review of Sensitivity and Workload Reduction'. Together they form a unique fingerprint.

Cite this