OSCA-Net: An Oriented-Structure Co-Attention Network for Vision-Based Structural Damage Detection
Main Article Content
Abstract
Visual inspection of concrete infrastructure is slow, subjective and hazardous, and the deep networks proposed to replace it are usually generic classifiers that ignore what a crack is. This paper presents OSCA-Net, a vision-based structural damage detector built around the geometry of surface distress rather than around a larger backbone. An oriented structure prior is computed without supervision from a multiscale Hessian ridge response modulated by structure-tensor coherence, so it responds to thin elongated dark features and suppresses isotropic staining. Pooled to the encoder grid, it enters every attention logit as a learnable per-head bias, guiding attention instead of cropping the image. Two frozen encoders, ResNet-50 and EfficientNet-B0, supply complementary token grids whitened to a common width by a projection fitted on training data only. Prior-biased co-attention blocks let each stream interrogate the other in both directions, and thirty-two photometric and Haralick descriptors gate the pooled visual code before classification. Only 1.04 million parameters are trained. On 40,000 concrete surface images OSCA-Net reaches 99.98 percent accuracy, 99.98 percent F1 and 0.9995 Matthews correlation, exceeding four published fine-tuned backbones on the same corpus. Clean accuracy there is saturated, every pretrained baseline falling within 1.15 points, so the models are separated instead by degradation: over four field corruptions at three severities OSCA-Net retains 98.88 percent against 97.00 percent for the weakest baseline.
