Grin logo
de en es fr
Shop
GRIN Website
Publish your texts - enjoy our full service for authors
Go to shop › Computer Sciences - Artificial Intelligence

Large Language Models in Scientific Research. A Systematic Review of Applications and Challenges

Title: Large Language Models in Scientific Research. A Systematic Review of Applications and Challenges

Research Paper (postgraduate) , 2026 , 24 Pages , Grade: N/A

Autor:in: Anonymous (Author)

Computer Sciences - Artificial Intelligence
Excerpt & Details   Look inside the ebook
Summary Excerpt Details

Large language models (LLMs) have emerged as a game-changer in computational tools and are expected to be used in hypothesis generation, drug discovery, literature analysis, medical question answering, and automated data processing. This systematic review aims to gather and combine extensive evidence on the use of LLMs in scientific research, together with the challenges encountered, according to the PRISMA 2020 guidelines. Systematic searches were undertaken in major academic databases from 2019–2025 for peer-reviewed studies and preprints. LLMs are applied to autonomous chemical research, multi-agent hypothesis generation systems, medical QA systems at expert level, and retrieval-augmented generation systems for knowledge synthesis. Major issues remain, including hallucinations and factual inaccuracy, reproducibility and methodological validation, privacy concerns and data security, and algorithmic bias. While LLMs can be highly useful and perform at an expert level in some scientific domains when properly validated and combined with domain knowledge, their deployment requires rigorous validation and field-specific fine-tuning. This review shows that strategic application of LLMs, particularly through self-reflection, knowledge graph integration, and retrieval augmentation, can significantly accelerate scientific discovery without compromising research integrity, while interdisciplinary cooperation is essential to develop governance frameworks, benchmarks, and ethical principles for sustainable scientific use.

Excerpt


Table of Contents

1. Introduction

2. Methodology

2.1 Review Protocol and Registration

2.2 Information Sources and Search Techniques

2.3 Inclusion and Exclusion Criteria

2.4 Screening and Study Selection

2.5 Data extraction and Quality Assessment

2.6 Data Synthesis and Analysis Approach

3. Results

3.1 Study Characteristics and Publication Landscape

3.2 Application in Literature Review and Knowledge Synthesis

3.3 Hypothesis Generation and Scientific Discovery

3.4 Applications to Drug Discovery and Molecular Design

3.5 Medical and Healthcare Applications

3.6 Applications of Materials Science and Chemistry

4. Significant problems and constraints.

4.1 Hallucination and Factual Inaccuracy

4.2 Reproducibility and Methodological Concerns

4.3 Privacy, Security and Data Protection

4.4 Bias and Fairness

5. Evidence-Based Solutions and Mitigation Strategies.

5.1 Self-Reflection and Verification Mechanisms

5.2 Domain-Specific Fine-Tuning and Adaptation

5.3 Retrieval-Augmented Generation and Knowledge Integration

6. Institutional Implementation And Governance Framework

6.1 Appropriate Use Guidelines

6.2 Transparency and Reproducibility Standards

6.3 Institutional LLM Literacy Programs

7. Discussion and Synthesis

7.1 Overall Evidence Quality and Strength

7.2 Critical Success Factors and Implementation Requirements

7.3 Limitations Of Current Evidence Base Gaps

8. Conclusion and Future Directions

Objectives & Topics

This systematic review investigates the empirical evidence surrounding the utilization of large language models across scientific research disciplines. By following the PRISMA 2020 guidelines, the study synthesizes the functional capabilities, practical applications, technical and methodological challenges, and evidence-based mitigation strategies associated with deploying generative AI in scientific inquiry, while defining governance recommendations for research institutions.

  • Empirical utility of large language models in literature synthesis, hypothesis generation, drug discovery, medical diagnostics, and materials science.
  • Major methodological and technical barriers, specifically hallucination phenomena, non-reproducible stochastic outputs, data privacy leakage, and systemic bias.
  • Technical mitigation strategies comprising retrieval-augmented generation, knowledge graph integration, self-reflection mechanisms, and low-rank adaptation.
  • Quality assessment and risk-of-bias evaluation of current peer-reviewed and preprint scientific literature on generative AI.
  • Institutional governance, author transparency benchmarks, and educational literacy frameworks for responsible AI deployment in academia.

Excerpt from the Book

3.2 Application in Literature Review and Knowledge Synthesis

The use of large language models has great potential for literature review and summarization of large scientific datasets. The comprehensive screening of literature and extraction of data from months to years are a bottleneck in evidence synthesis as is done in traditional systematic reviews. These timeframes are significantly reduced with LLM-based approaches.

Standard LLMs suffer from a number of basic shortcomings, which can be mitigated by retrieval-augmented generation (RAG) systems that combine knowledge retrieval and generative capabilities (Asai et al., 2024; Lewis et al., 2020). The information retrieved from the documents or knowledge snippets can be used to fine-tune prompts for LLM generation, unlike hardcoded parametric knowledge. This greatly reduces hallucinations and increases factual accuracy. Brown et al. (2025) conducted a systematic review of RAG techniques, summarizing architectures, evaluation metrics and identifying limitations. In scientific document structured information extraction, RAG systems have been found to be more efficient than traditional LLM-based methods, achieving 85-92% accuracy compared to 60-75%.

Moreover, integrating knowledge graphs helps to enhance the accuracy by restricting the output of LLMs to known entities and relationships in structured knowledge representations. Gao et al. (2023) and Edge et al. (2024) present graph-based RAG systems, which build a knowledge graph from a collection of documents, and use LLMs to retrieve subgraph structures and perform multi-hop reasoning. These systems were especially successful with complex queries that involved synthesizing information from several documents.

However, challenges remain. 15-25% of the published literature syntheses hallucinate non-existent references, particularly for less common publications or new research domains where training data is not well-represented. Temporal aspects are challenging—LLM knowledge is limited to data distributions, which can result in gaps in knowledge for very recent publications that have come out after model training. Difficult to incorporate collections of documents that are continually revised.

Chapter Summaries

1. Introduction: Outlines the emergence of transformer-based large language models in scientific research, highlighting their transition from traditional supervised machine learning to general-purpose context-aware engines alongside emerging operational risks.

2. Methodology: Details the systematic review protocol adhering to PRISMA 2020 and prospective PROSPERO registration, delineating database search strategies, inclusion and exclusion criteria, dual-reviewer screening, and methodological quality scoring across 45 analyzed studies.

3. Results: Categorizes evidence across seven scientific application domains, assessing the quantitative performance of LLMs in literature extraction, autonomous hypothesis formulation, molecular design, medical reasoning, and materials engineering.

4. Significant problems and constraints.: Critically analyzes fundamental failure modes, including variable hallucination rates, stochastic output variance threatening scientific reproducibility, data memorization breaches under privacy statutes, and demographic or geographic biases.

5. Evidence-Based Solutions and Mitigation Strategies.: Evaluates architectural and algorithmic solutions designed to bolster LLM reliability, focusing on Self-RAG frameworks, knowledge graph integrations, parameter-efficient fine-tuning via LoRA, and external verification modules.

6. Institutional Implementation And Governance Framework: Formulates actionable operational standards for academic institutions, emphasizing hybrid human-in-the-loop workflows, mandatory peer-review transparency disclosures, and institutional AI literacy training.

7. Discussion and Synthesis: Weighs the overall quality of available evidence, contrasts critical implementation requirements with pervasive methodological deficits across current literature, and outlines limitations inherent to preprint-heavy AI research.

8. Conclusion and Future Directions: Reasserts the value of LLMs as assistive augmentative tools rather than autonomous authorities, calling for interdisciplinary cooperation between domain scientists, computer scientists, and ethicists to safeguard research integrity.

Keywords

Large language models, Systematic review, PRISMA guidelines, Hallucination, Retrieval-augmented generation, Hypothesis generation, Drug discovery, Reproducibility, Knowledge extraction, AI governance, Domain-specific fine-tuning, Scientific methodology

Frequently Asked Questions

What is the core subject of this systematic review?

The work provides an exhaustive, multidisciplinary evaluation of how large language models are employed across diverse scientific disciplines, examining both their empirical performance gains and the methodological risks they pose to scientific integrity.

Which scientific disciplines are primarily explored?

The review examines applications across biomedical research, pharmacology and drug discovery, clinical medicine, chemistry and materials science, data science, computational programming, and systematic literature synthesis.

What are the primary objectives and research questions of the study?

The study investigates the documented evidence of LLM efficacy in scientific tasks, the technical and ethical barriers hindering reliable deployment, verified mitigation approaches to overcome limitations, and policy structures required for academic institutions.

What scientific methodology does the review adhere to?

The authors followed the PRISMA 2020 guidelines for systematic reviews and prospectively registered the protocol in PROSPERO, conducting systematic searches across six multidisciplinary databases (Scopus, Web of Science, IEEE Xplore, ACM Digital Library, ScienceDirect, and SpringerLink) between 2019 and 2025.

What topics are analyzed within the main results and constraints chapters?

The main body examines empirical tasks such as autonomous chemical experimentation, clinical licensing exam performance, and molecular property prediction, while contrasting these achievements against hallucination rates, unrecorded hyperparameters, data privacy vulnerabilities, and demographic performance discrepancies.

What keywords best characterize the publication?

Key terms include large language models, systematic review, PRISMA guidelines, hallucination, retrieval-augmented generation, hypothesis generation, drug discovery, reproducibility, and AI governance.

How widespread is the hallucination phenomenon across different scientific tasks?

The review notes that hallucination rates vary widely by domain complexity, ranging from 15–25% in structured literature synthesis to 35–45% in open-ended scientific hypothesis generation, and peaking at 50–82% when models query specialized or newly emerging research topics.

Why do traditional LLM-generated hypotheses often fail experimental validation?

While LLMs achieve high plausibility ratings (60–75%) from expert reviewers during blind evaluations, only 15–35% withstand laboratory testing. This divergence occurs because models excel at producing rhetorically fluent associations matching training patterns rather than mechanistically sound physical, chemical, or biological pathways.

How do Retrieval-Augmented Generation (RAG) and knowledge graphs resolve standard LLM limitations?

RAG systems restrict generative outputs by grounding prompts in retrieved factual source documents rather than relying strictly on internal parametric weights, raising information extraction accuracy to 85–92%. Integrating structured knowledge graphs further enables verified multi-hop reasoning across multi-document sets while establishing clear data provenance.

What governance and transparency standards do the authors propose for academic institutions?

The authors advocate for mandatory disclosure statements regarding LLM utilization in academic publishing, comprehensive documentation of model parameters, seeds, and prompts to preserve reproducibility, and the retention of final decision-making power strictly in human expert hands.

Excerpt out of 24 pages  - scroll top

Details

Title
Large Language Models in Scientific Research. A Systematic Review of Applications and Challenges
College
The Islamia University of Bahawalpur  (The Islamia University of Bahawalpur)
Course
Computer Science
Grade
N/A
Author
Anonymous (Author)
Publication Year
2026
Pages
24
Catalog Number
V1759642
ISBN (PDF)
9783389204931
ISBN (Book)
9783389204948
Language
English
Tags
Large language models Scientific research Systematic review Hallucinations Drug discovery
Product Safety
GRIN Publishing GmbH
Quote paper
Anonymous (Author), 2026, Large Language Models in Scientific Research. A Systematic Review of Applications and Challenges, Munich, GRIN Verlag, https://www.grin.com/document/1759642
Look inside the ebook
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
  • Depending on your browser, you might see this message in place of the failed image.
Excerpt from  24  pages
Grin logo
  • Grin.com
  • Shipping
  • Contact
  • Privacy
  • Terms
  • Imprint
  • Withdraw Contract