Gottweis et al. 2026 show that scaling test-time compute—through self-play debate, hypothesis tournaments and an evolution process—continues to improve hypothesis quality over time. They validate Co-Scientist in three biomedical settings of varied complexity: cancer drug repurposing, liver-fibrosis target discovery, and antimicrobial-resistance mechanism identification.
E5 confirms test-time compute scaling with continued hypothesis-quality improvements, and E6 confirms the self-play debate, hypothesis tournaments, evolution process, and the three biomedical validation areas (cancer drug repurposing, liver-fibrosis target discovery, AMR mechanism identification) exactly as claimed.
Written by Kimi K3 via Ollama Cloud · checked by GLM-5.3 via Ollama Cloud · 30 Sept, 23:08
Source chain
AEvery quote below was checked, without a model, to appear verbatim in its source.
- 01
“Automated evaluations show continued benefits of test-time compute scaling, improving hypothesis quality over time.”
Full-text passage · no page number · evidence E5
Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.
- 02
“drug repurposing for cancer; novel treatment target discovery for liver fibrosis; and identification of mechanistic explanations for antimicrobial resistance”
Full-text passage · no page number · evidence E6
Passage read from www.ebi.ac.uk, which may be a preprint rather than the published version.