The Benchmark Beneath The BenchMark
A model can score well against a test set and still be unsafe for the people it is meant to serve. Health AI benchmarking is necessary, but incomplete: if the evidence beneath the score does not represent the intended population, the model is only reproducing that gap.
This briefing sets out why representation must be inspectable before a performance score can be trusted. It introduces SökerData’s twelve-dimension data benchmark, six questions about who is included, and six about how evidence is produced and used, and argues that participation, retention, bandwidth and care pathway shape what any later model test can honestly claim.
It is offered for discussion, not as a field-wide standard.