Event date · · arXiv

Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations

FACT STATEMENT

A review paper synthesizes 57 method-centered papers on class activation mapping (CAM) published from 2016 onward. The paper develops a taxonomy separating methods by attribution mechanism, architectural dependence, and evaluation objective. It reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. The paper notes the field is shifting from CNN-specific methods to transformer and foundation-model-era approaches.

What happened

Class activation mapping (CAM) is a widely used visual explanation family in explainable AI that converts internal model evidence into heatmaps highlighting image regions, channels, tokens, or patches supporting a target class. Since the first CAM formulation in 2016, the field has expanded beyond global-average-pooled CNN classifiers to include gradient-based post-hoc explanations, gradient-free score and ablation methods, high-resolution upscaling, weakly supervised localization and segmentation, transformer token attribution, causal and debiasing methods, and foundation-model-era approaches using CLIP, DINO, SAM, or feature-distribution comparisons. This review synthesizes a strict corpus of 57 method-centered papers published from 2016 onward, developing a taxonomy that separates methods by attribution mechanism, architectural dependence, and evaluation objective. It reviews gradient-based CAMs, recent and hybrid CAM-style methods, and model-based or architecture-aware methods. The main trend is a shift from CNN-specific methods to transformer and foundation-model-era approaches.

Technical significance

The review indicates that CAM methods are evolving from gradient-based post-hoc explanations tied to CNN architectures toward gradient-free, ablation-based, and architecture-aware approaches that can handle transformer token attributions and foundation-model feature distributions. This suggests a need for evaluation frameworks that account for architectural dependence and attribution mechanism.

Industry impact

As visual explanation methods expand to foundation models like CLIP, DINO, and SAM, industries relying on computer vision for high-stakes decisions may need to update their explainability tooling and validation practices to accommodate new attribution mechanisms and model architectures.

Decision value

The review provides a structured taxonomy and corpus for practitioners selecting visual explanation methods, potentially reducing evaluation overhead and guiding adoption of CAM variants suited to modern architectures.

What to watch

Observable next signals include new CAM-style methods targeting transformer and foundation-model architectures, standardized benchmarks for evaluating explanation quality across architectures, and increased adoption of gradient-free or ablation-based attribution in production systems.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.