Evaluation and methods for domain-grade models.
Open benchmarks, technical writing, and evaluation frameworks from the Zstate team. Agent trajectories, clinical reasoning, and regulated-domain preference data.
- RxScribe-Bench: Structured Extraction from Indian Outpatient Prescriptions, with Hallucination and Abstention as First-Class Safety MetricsResults across four axes, and what happened when two of them turned out to be paying models for not answering.
- The Missing Benchmark: Why AI's Next Crisis Will Be About Data, Not Models
No research entries match your search.