OpenHub Repository

Comparative evaluation of parallel and sequential hybrid CNN–ViT models for wrist X-ray anomaly detection

Loading...
Thumbnail Image

Date

Journal Title

Journal ISSN

Volume Title

Publisher

Sol Plaatje University

Abstract

Medical anomaly detection is often challenged by limited annotated data and domain shifts, which constrain the performance and generalization of Deep Learning (DL) models. Hybrid Convolutional Neural Network–Vision Transformer (CNN–ViT) architectures have shown strong potential, yet they typically rely on large, welllabeled datasets. The Convolutional Neural Network (CNN) captures local spatial features, while the Vision Transformer (ViT) models global contextual relationships, making their combination particularly powerful. Multistage Transfer Learning (MTL) offers a practical strategy to mitigate data scarcity and domain variation. In this study, we evaluated two CNN–ViT fusion strategies: parallel, where CNN and ViT features are extracted independently and fused before classification, and sequential, where CNN features are passed through the ViT for integrated processing. Models were pretrained on non-wrist musculoskeletal radiographs (MURA), fine-tuned on the MURA wrist subset, and evaluated for cross-domain generalization using an external wrist X-ray dataset from the Al-Huda Digital Xray Laboratory. The Xception–DeiT(Data-efficient Image Transformer) parallel hybrid achieved the strongest internal performance (accuracy = 0.88), whereas the sequential DenseNet–ViT generalized best in zero-shot transfer. After light finetuning, parallel hybrids achieved near-perfect accuracy (0.98) and recall (1.00). Statistical analyses showed no significant difference between the parallel and sequential models (McNemar’s test), while backbone selection played a key role in performance. The Wilcoxon test revealed no significant difference in recall and F1-score between image and patient-level evaluations, indicating balanced performance across both levels. Sequential hybrids achieved up to sevenfold faster inference than parallel models on the MURA test set while maintaining similar Graphics Processing Unit (GPU) memory usage (3.7 GB). Both architectures produced clinically meaningful saliency maps highlighting relevant wrist regions. Overall, this study provides a systematic comparison of CNN–ViT fusion strategies for wrist anomaly detection, clarifying trade-offs between accuracy, generalization, interpretability, and computational efficiency in clinical artificial intelligence (AI).

Description

Keywords

Citation

Endorsement

Review

Supplemented By

Referenced By

Creative Commons license

Except where otherwised noted, this item's license is described as Attribution-NonCommercial-NoDerivs 3.0 United States