Claim 60 · ai-for-science · connection

Despite the autonomy claims, current models show sharp reasoning limits that keep humans in the loop: Alampara et al. 2025 find vision-language models fail at spatial reasoning, cross-modal synthesis and multi-step inference, concluding they cannot yet serve as autonomous scientific reasoners. Consistent with this, Minasny et al. 2026 report that LLMs answered only up to 65% of advanced soil-science examination questions.

Supported

E40/E39 report fundamental limitations in spatial reasoning, cross-modal synthesis and multi-step inference and conclude VLLMs cannot yet serve as autonomous scientific reasoners, and E29 reports LLMs correctly answering only up to 65% of advanced soil-science exam questions with human validation still required.

Written by Kimi K3 via Ollama Cloud · checked by GLM-5.3 via Ollama Cloud · 30 Sept, 23:08

Source chain

A

Every quote below was checked, without a model, to appear verbatim in its source.

  1. 01

    “fundamental limitations in spatial reasoning, cross-modal information synthesis and multi-step logical inference”

    Full-text passage · no page number · evidence E40

    Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.

  2. 02

    “current VLLMs cannot yet serve as autonomous scientific reasoners”

    Full-text passage · no page number · evidence E39

    Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.

  3. 03

    “they could correctly answer only up to 65% of questions from advanced soil science examinations”

    Full-text passage · p. 3 · evidence E29

    While they demonstrate promising capabilities for conversational interaction, ensuring their reliability and depth of understanding in specialized soil science domains still requires signifi cant human validation and expertise. For example, Khanifar ( 34) assessed the performance of LLMs in answering soil science-related questions and found that, on average, they could correctly answer only up to 65% of questions from advanced soil science examinations. This existing foundation, with its strengths and weaknesses, sets the stage for the development of next-generation AI agents capable of more integrated, autonomou s, and collaborative scienti fic exploration and management in soil science. The vision of AI agents transforming scienti fic discovery is gaining momentum in fields such as chemistry, materials science, and biomedicine, where complex systems, vast datasets, and multidisciplinary knowledge must be integrated ( 5, 35, 36). The concept of an “AI scientist” or “co-scientist”— an AI system capable of skeptical learning, reasoning, and collaboration— has been proposed as a framework for generating novel scienti fich y p o t h e s e s aligned with researcher objectives ( 8, 37, 38).

    Passage read from www.frontiersin.org, which may be a preprint rather than the published version.