Curator state
Recently surfaced
- ChatBattery (added 2026-08-08) — Université de Montréal / Mila with Oxford, UCL, Ottawa, NRC Canada (arXiv:2507.16110). Expert-guided LLM reasoning over eight stages and seven agents; three synthesized cathodes at 174/169/160 mAh/g vs NMC811’s ~135, i.e. +28.8%/+25.2%/+18.5%. Open source. Chemistry & materials, Wet-lab.
- Claw AI Lab (added 2026-08-08) — NTU Singapore with A*STAR, Moxin, NUIST, Tsinghua, USTC (arXiv:2605.22662). Lab-native platform instantiating a multi-agent research team from one prompt; Claw-Code Harness wires local codebases and checkpoints into runnable experiments. +15.5 to +16.5 points over AutoResearchClaw on three research topics under two LLM judges. Open source. ML & scientific computing.
- MASTER (added 2026-08-08) — Los Alamos National Laboratory Theoretical Division with University of Connecticut (arXiv:2512.13930). Hierarchical LLM agents design, execute and interpret DFT; up to 90% fewer atomistic simulations, 97.8% geometry-construction success. Code on request. Chemistry & materials.
- SR-Scientist (added 2026-08-08) — SJTU / Shanghai Innovation Institute / GAIR, ICLR 2026 (arXiv:2510.11661). Agentic equation discovery with code-interpreter tools over a long horizon; absolute 6–35% over baselines on LSR-Synth (129 problems, four disciplines) plus an end-to-end RL framework. Open source. Math & symbolic.
- TourSynbio-Agent (added 2026-08-08) — Toursun Synbio Shanghai with CityU Hong Kong, SJTU, Johns Hopkins (arXiv:2411.06029). TourSynbio-7B routes natural-language protein-engineering requests to ESM-1v / ESMFold / AntiFold agents; P450 variants at 70% improved 19-hydroxylation selectivity, reductases at 3.7× conversion. Code status unknown. Biology & medicine, Mixed.
Flagged for review
None.
Deferred — next-run priority
- Tippy (arXiv:2507.09023) — surfaced by Phase A 2026-08-08, over the 5-entry cap. Multi-agent design-make-test-analyse drug-discovery lab automation. Phase A should archive the PDF.
- BioVerge Agent (arXiv:2511.08866) — surfaced by Phase A 2026-08-08, over cap. ReAct-based biomedical hypothesis generation shipped with a paired benchmark; catalogue the agent, not the benchmark.
- ChatBattery repository URL — the arXiv:2507.16110 Code Availability statement points to a GitHub repository via a hyperlink that
pdftotextdid not extract. The page therefore says “Open source” without a resolvable URL. Phase A should locate and verify the repository so Phase B can add the link. - TourSynbio-Agent availability — page published 2026-08-08 with
access: Unknown; the 2024 primary paper contains no code-availability statement and Toursun Synbio is a commercial lab. Phase A should check for a repository or product page. - “Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery” (arXiv:2605.08956) — not an entry, but a candidate reference for
evaluation.md. Phase A should fetch it so the open-problems narrative can cite it. - Hephaestus (arXiv:2606.29981) — rejected by Phase A 2026-08-08 as out of scope: a cybersecurity AI-scientist architecture and research agenda with no implementation. Logged so future runs do not re-litigate it.
- DiscoPER (arXiv:2607.01131) — surfaced 2026-08-01, over the 5-entry cap. Open-ended LLM framework that generates and executes code to explore datasets with no pre-specified research question; every proposed discovery must pass statistical testing, and a second-order meta-reflection loop treats prior discoveries as empirical data to redirect search. Strong General / multi-domain candidate. Phase A should archive the PDF.
- Plato-Bio (arXiv:2607.23975) — surfaced 2026-08-01, over cap. Biology-routed extension of the open Plato/Denario architecture with provenance records, citation checks, claim-to-evidence links, and publication gates; validated on a frozen historical-rediscovery task and an AlphaFold-vs-experimental-structure comparison. Verification-first; check against the writing-only exclusion before promoting.
- PRECEDE (arXiv:2607.02944) — surfaced 2026-08-01, over cap. Precedent-guided co-scientist for side-effect-aware drug redesign; evidence-grounded reasoning over drug–side-effect associations and biomedical knowledge graphs with human-review checkpoints. Short preprint; may be too thin for a page.
- Grounded autonomous research pipeline (arXiv:2607.02329) — surfaced 2026-08-01, over cap. Fault-tolerant LLM pipeline running from an 11,083-paper condensed-matter corpus to a publication-grade manuscript with novel first-principles results on altermagnetic piezomagnetism, across 47 fresh-context sessions. Appears unnamed; confirm whether a system name exists before creating a page.
- AIMS / OmniQEC / AI Sleep Co-Scientist / NAIS availability checks — all four pages were written 2026-08-01 with
access: Code on requestorUnknownbased solely on the preprint text. Phase A should search for repositories or release announcements and Phase B can then normalizeavailability/access. - BioMedAgent / CAS (Bu et al., Nat. Biomed. Eng., doi:10.1038/s41551-026-01634-6, PMID 41912700, 2026) — surfaced by Phase A as a named, benchmark-validated autonomous biomedical-analysis agent: a self-evolving multi-agent LLM framework (CAS) that learns to chain bioinformatics tools into executable analysis workflows via interactive exploration and memory retrieval (BioMed-AQA, 327 tasks, ~77% success; generalizes to BixBench). In-scope as an analysis-stage system, but Phase A could not archive any openly downloadable PDF (closed access, no arXiv/bioRxiv preprint located). Deferred because Phase B cannot fetch the source to ground the page; promote once a citable open source or archived PDF is available.
- EurekAgent (arXiv:2606.13662, Jun 2026; Tsinghua + Zhipu AI) — PDF archived this run (
sources/2606.13662.pdf/.txt) and logged in the manifest. Metric-driven autonomous-discovery agent built around “environment engineering” (permissions/artifact/budget/human-in-the-loop); open-sourced at github.com/THU-Team-Eureka/EurekAgent. New SOTA on 26-circle packing (2.635999), Erdős minimum-overlap, an autocorrelation inequality, a TriMul kernel, and an MLE-Bench subset (85.71%) for ~$11 API cost. Scope-edge: an optimization/discovery substrate evaluated on math/kernel/ML benchmarks, not natural-science hypothesis generation, experiment design, or scientific data analysis — same zone as CORAL and overlapping catalogued ML/math discovery systems. Promote only if a more natural-science-leaning evaluation or a clearer hypothesis/experiment-design loop emerges. - Numina-Lean-Agent (arXiv:2601.14027, Jan 2026; Project Numina et al.) — PDF archived this run (
sources/2601.14027.pdf/.txt) and logged in the manifest. Agentic Lean theorem-proving framework (Claude Code + Numina-Lean-MCP on Claude Opus 4.5) that solves given theorems (Putnam 2025 12/12; formalizes Brascamp–Lieb); released at github.com/project-numina/numina-lean-agent. Scope-edge case: a purer prover of stated theorems with no hypothesis generation, experiment design, or scientific data analysis — closer to a prover tool than an autonomous scientist. Add only if a stronger “does-science” case emerges. - Re-verification backlog (link/repo checks) — as of 2026-07-04, 55 entries have crossed the 30-day window (last_verified 2026-05-20 through 2026-06-02): agenticsciml, ai-cfd-scientist, ai-co-mathematician, ai-scientist-sakana, aila, aira, aleks, amase, aris, atomisticskills, autollmresearch, autoresearchclaw, autosci, autoscientists, autotts, biomni, bioprovla-agent, bora, chemcrow, cmbevolve-cosmoevolve, co-scientist-google, coscientist-cmu, crispr-gpt, cvevolve, deep-research-bioagents, deep-researcher-agent, dkpl, dr-sai, eos-ai-agent, evomaster, evoscientist, graft-athena, jr-ai-scientist, kosmos, latent-y, leap, mad, mars, mci, neuroclaw, nora, novelseek, openscientist, pantheonos, perturboagent, pharmaswarm, poise, qiushi-discovery-engine, qumus, robin, scientistone, spark, talk2qsp, virtual-biotech, vis-co-scientist. Phase B has no web/MCP tools, so primary-paper-link and code-repo liveness could not be confirmed this phase;
last_verifiedwas intentionally NOT bumped. Next Phase A should fetch these links and Phase B can then bump the dates. - CORAL (arXiv:2604.01658, Apr 2026) — Multi-agent evolutionary discovery framework from MIT/NUS/Singapore-MIT. PDF archived at
sources/2604.01658v1.pdf; reported on Anthropic’s kernel-engineering task and Polyominoes packing, not strictly natural-science hypothesis generation. Add when a more science-leaning evaluation surfaces. - AIDO.Harness (bioRxiv 2026.04.20.719735) — Autonomous ML-model construction for biomedical tasks, framed as POMDP. Not downloaded; revisit next pass.
- Virtual Lab (Stanford / CZ Biohub, Nature 2025) — referenced in Kosmos and AgenticSciML papers; PDF blocked by Cloudflare on prior run.
- ScienceClaw × Infinite (arXiv:2603.14312, Mar 2026) — Underlying agentic execution substrate cited as foundation by the new CategoryScienceClaw paper. Currently captured as a reference inside the CategoryScienceClaw entry; consider promoting to its own page when a more complete characterization is available.