mai₂ lab.

Explaining Vision Models via Game-Theoretic Indices

Explaining Vision Models via Game-Theoretic Indices

Deep vision models achieve high accuracy, but their decisions are often a black box. We use game-theoretic frameworks — Shapley values and related notions — to quantitatively analyze how different regions of an image (pixels or patches) jointly contribute to a model’s prediction.

Rather than scoring individual pixels in isolation, we treat a model’s decision as a “knowledge structure” that emerges from interactions between regions. We’re also investigating a subtler issue: averaging-based attributions like the Shapley value can systematically over-credit spurious “contextual distractor” groups depending on context. To address this, we study more robust, stability-based attribution methods grounded in concepts like the least-core.

Progress so far

We presented work on identifying important groups of pixels through game-theoretic interactions at CVPR 2024, and extended this to object detectors by explaining their predictions through the collective contribution of pixels. We’ve also formalized how Shapley values can over-attribute importance to contextual distractors, and proposed a more stable attribution method based on the least-core (ICML 2026 workshop).

Related Publications

* Corresponding author

  • Seeing Through Distractions: Stable Attribution via the Core

    Sai Ganesh Nagarajan, Toshinori Yamauchi, Hiroshi Kera

    International Conference on Machine Learning (ICML) Workshop, 2026

  • Explaining Object Detectors via Collective Contribution of Pixels

    Toshinori Yamauchi, Hiroshi Kera, Kazuhiko Kawamoto

    Meeting on Image Recognition and Understanding (MIRU 2025), 2025

  • Identifying Important Group of Pixels using Interactions

    Kosuke Sumiyasu, Kazuhiko Kawamoto, Hiroshi Kera*

    Meeting on Image Recognition and Understanding (MIRU 2024), 2024