Cura 1T: Specialized Model for Agentic Healthcare
Abstract
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.
Community
Cura 1T leads frontier models on 5 of 6 hardest healthcare benchmarks:
- HealthBench Hard: 36.8 (GPT-5.5: 31.5)
- HealthBench Professional: 66.2 (Claude Fable 5: 66.0)
- MedXpertQA-Text: 60.0 (GPT-5.5: 59.6)
- MedXpertQA-Multimodalt: 72.2 (GPT-5.5: 77.1)
- AgentClinic: 79.6 (Claude Opus 4.8: 79.4)
- MedAgentBench-v2: 94.0 (Claude Opus 4.8: 93.7)
How we trained it: RSI (recursive self-improvement).
Each iteration, a training agent plans a target capability, trains the model, evaluates the graded benchmark trajectories, and data agent synthesizes the next data mixture from the failure modes it finds.
Humans gate every keep-or-revert decision. Reverted rounds stay in the record. One raised headline scores while quietly damaging a held-out subset, so we threw it out. The kept rounds add up: +14.6 on HealthBench Hard, +15.9 on HealthBench Professional, +9.3 on MedAgentBench.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AutoMedBench: Towards Medical AutoResearch with Agentic AI Models (2026)
- Towards Autonomous and Auditable Medical Imaging Model Development (2026)
- HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents (2026)
- Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance (2026)
- EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning (2026)
- MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Models for Clinical Reasoning (2026)
- CAREAgent: Clinical Agent with Structured Reasoning and Tool-Integrated for Order Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.15314 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
