Paper-clean v2 currently has smoke-test results only. Full v2 evaluation should be run with the model-specific manifests in
benchmark_v2_model_inputs/by_model.Collected CT-volume, video-slice, and 2D proxy benchmark results for the current radiology VLM work. Metrics are shown within each benchmark family only; native 3D, video-slice, and 2D proxy inputs should not be merged into one leaderboard.
benchmark_v2_model_inputs/by_model.