Maximize your thought leadership

Diffusion-Mamba Hybrid Achieves Breakthrough in Infrared Small Target Detection

By Burstable Editorial Team•
Researchers at Harbin Institute of Technology have developed a two-stage deep learning network that combines diffusion-based feature enhancement with state-space modeling to achieve state-of-the-art infrared small target detection, with significant implications for remote sensing, wildfire prevention, and surveillance.
Diffusion-Mamba Hybrid Achieves Breakthrough in Infrared Small Target Detection

Infrared small target detection is crucial for applications ranging from forest fire early warning to remote sensing threat assessment, but existing methods struggle with tiny, low-contrast targets that lack distinct shape or texture. A new approach developed by researchers at the Research Center for Space Optical Engineering at Harbin Institute of Technology, published on June 30, 2026, in the Journal of Remote Sensing, promises to set a new standard for this challenging task.

The proposed Diffusion-Enhanced Dense Mamba Network (DEDM-Net) addresses a critical challenge: infrared targets often occupy fewer than 81 pixels (typically under 9×9) and exhibit extremely low energy with signal-to-noise ratios around 3, causing them to be easily submerged in background clutter. Most deep learning methods focus exclusively on target features while neglecting background information, leading to severe class imbalance. The DEDM-Net overcomes this by simultaneously modeling both targets and backgrounds.

The two-stage network achieves a synergistic effect. The first stage employs a dual-path diffusion model with a novel blind processing module that predicts each pixel using only surrounding information—never the pixel itself—preventing extremely small targets from being misclassified as background. The second stage introduces a dense nested Mamba architecture based on state-space modeling (SSM), which captures long-range correlations across global and local features with linear computational complexity, a significant advantage over conventional Transformers. A cross-stage prediction fusion module further integrates features from both stages, improving contour segmentation accuracy.

Evaluated on three public datasets—NUAA-SIRST, NUDT-SIRST, and IRSTD-1k—the DEDM-Net outperformed 11 state-of-the-art methods. On NUDT-SIRST, it achieved 93.40% IoU, 93.28% nIoU, 98.37% detection probability (P_d), and a false-alarm rate of just 3.75×10⁻⁶. On IRSTD-1k, it reached 73.71% IoU and 93.89% P_d with only 11.10×10⁻⁶ false alarms. These results demonstrate significant reductions in both missed detections and false alarms in complex imaging environments.

"Infrared small targets are extremely challenging because they lack shape and texture—they're essentially just a few bright pixels in a sea of background," said corresponding author Dr. Shikai Jiang. "By modeling both the target-free background and potential target regions simultaneously, our diffusion-enhanced approach effectively amplifies what matters while suppressing what doesn't. The Mamba architecture then provides the global context needed to distinguish true targets from bright clutter."

The method employs a two-stage training framework. In the diffusion enhancement stage, a U-Net backbone estimates noise across 1,000 diffusion steps, with a dual-path scheme modeling both target masks and background images. The blind processing module generates pixel-wise convolution kernels that exclude the center pixel, effectively removing small targets from background reconstruction. In the detection stage, a dense nested network with feature pyramid connections and residual Mamba blocks extracts multiscale features. The loss function combines binary cross-entropy and Dice losses to address class imbalance.

While DEDM-Net achieves superior accuracy, the diffusion-based two-stage design increases inference time compared to single-stage networks. Future work will focus on model distillation, mixed-precision inference, and faster samplers to reduce the required diffusion steps. The approach holds promise for real-time surveillance systems, autonomous drone navigation in low-visibility conditions, and early wildfire detection networks. As the team noted, the framework could also inspire new thinking about how generative models and state-space architectures can be combined for other challenging computer vision tasks where target-background separation is critical.

The research was supported by the National Natural Science Foundation of China, the China Postdoctoral Science Foundation, the Natural Science Foundation of Heilongjiang Province, and the National Key Laboratory of Air-Based Information Perception and Fusion. The full study is available at https://spj.science.org/doi/10.34133/remotesensing.1046, and additional information can be found at http://chuanlink-innovations.com.

Burstable Editorial Team

Burstable Editorial Team

@burstable

Burstable News™ is a hosted solution designed to help businesses build an audience and enhance their AIO and SEO press release strategies by automatically providing fresh, unique, and brand-aligned business news content. It eliminates the overhead of engineering, maintenance, and content creation, offering an easy, no-developer-needed implementation that works on any website. The service focuses on boosting site authority with vertically-aligned stories that are guaranteed unique and compliant with Google's E-E-A-T guidelines to keep your site dynamic and engaging.