Evaluation
Long-form references on benchmarks, measurement, and what evaluations actually test.
Research on safe, trustworthy, and verifiable AI systems.
Long-form references on benchmarks, measurement, and what evaluations actually test.
Contributed to an open-source benchmark suite for medical LLM capabilities.