TY - JOUR
T1 - Controllable FrFT-Driven Y-Net for Cross-Modal Image Fusion Scenarios
AU - Jiang, Yifan
AU - Li, Kemin
AU - Li, Wenyuan
AU - Huang, Guoheng
AU - Yang, Jietao
AU - Yuan, Xiaochen
AU - Chen, Xuhang
AU - Ling, Bingo Wing Kuen
AU - Pun, Chi Man
AU - Wang, Shunlan
N1 - Publisher Copyright:
© 1975-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - Image fusion aims to synthesize high-quality representations by integrating complementary information from multimodal sources. While recent deep learning approaches have explored spectral domain learning to overcome the localized receptive fields of spatial convolutions, they predominantly rely on the standard Fourier Transform. However, two critical bottlenecks remain: standard Fourier bases assume signal stationarity (ill-suited for non-stationary natural images) and lack adaptive frequency-adaptive feature weighting, leading to spectral leakage, ineffective high-frequency-noise decoupling, and texture smoothing. To break these limitations, we propose the Fractional-Order Dynamic Perceptive Y-Net (FDP Y-Net), a unified framework driven by Controllable Fractional Fourier Transform (FrFT) Convolutions. By generalizing the spectral domain, our model adaptively rotates the time-frequency axis to optimally represent non-stationary features and dynamically weights frequency components. Specifically, a Dynamic Perceptive Module spatially localizes salient regions, the Controllable FrFT Block captures time-varying spectral characteristics, and an Information Enhancement Module explicitly reconstructs high-frequency residuals. Orchestrated within a Y-shaped architecture with specialized skip connections, these components ensure holistic feature fusion. Extensive experiments across Infrared-Visible, Medical, and Multifocus datasets demonstrate that FDP Y-Net effectively resolves spectral limitations and frequency weighting deficiencies, achieving state-of-the-art performance in visual fidelity and quantitative metrics.
AB - Image fusion aims to synthesize high-quality representations by integrating complementary information from multimodal sources. While recent deep learning approaches have explored spectral domain learning to overcome the localized receptive fields of spatial convolutions, they predominantly rely on the standard Fourier Transform. However, two critical bottlenecks remain: standard Fourier bases assume signal stationarity (ill-suited for non-stationary natural images) and lack adaptive frequency-adaptive feature weighting, leading to spectral leakage, ineffective high-frequency-noise decoupling, and texture smoothing. To break these limitations, we propose the Fractional-Order Dynamic Perceptive Y-Net (FDP Y-Net), a unified framework driven by Controllable Fractional Fourier Transform (FrFT) Convolutions. By generalizing the spectral domain, our model adaptively rotates the time-frequency axis to optimally represent non-stationary features and dynamically weights frequency components. Specifically, a Dynamic Perceptive Module spatially localizes salient regions, the Controllable FrFT Block captures time-varying spectral characteristics, and an Information Enhancement Module explicitly reconstructs high-frequency residuals. Orchestrated within a Y-shaped architecture with specialized skip connections, these components ensure holistic feature fusion. Extensive experiments across Infrared-Visible, Medical, and Multifocus datasets demonstrate that FDP Y-Net effectively resolves spectral limitations and frequency weighting deficiencies, achieving state-of-the-art performance in visual fidelity and quantitative metrics.
KW - Dynamic Perception
KW - Fractional Fourier Transform (FrFT)
KW - High-frequency Enhancement
KW - Image Fusion
KW - Non-stationary Signal Processing
UR - https://www.scopus.com/pages/publications/105039599570
U2 - 10.1109/TCE.2026.3694692
DO - 10.1109/TCE.2026.3694692
M3 - Article
AN - SCOPUS:105039599570
SN - 0098-3063
JO - IEEE Transactions on Consumer Electronics
JF - IEEE Transactions on Consumer Electronics
ER -