Skip to main content

Motion-temporal calibration network for continuous sign language recognition.

Hu, H., Peng, J., Xiao, Z., Guo, L. and Wu, D., 2025. Motion-temporal calibration network for continuous sign language recognition. Complex and Intelligent Systems, 12, 35 (2026).

Full text available as:

[thumbnail of OPEN ACCESS]
Preview
PDF (OPEN ACCESS)
s40747-025-02156-5.pdf - Published Version
Available under License Creative Commons Attribution.

2MB
[thumbnail of Title Page.pdf] PDF
Title Page.pdf - Accepted Version
Restricted to Repository staff only

151kB
[thumbnail of MTCNet.pdf] PDF
MTCNet.pdf - Accepted Version
Restricted to Repository staff only

1MB

DOI: 10.1007/s40747-025-02156-5

Abstract

Continuous Sign Language Recognition (CSLR) is fundamental to bridging the communication gap between hearing-impaired individuals and the broader society. The primary challenge lies in effectively modeling the complex spatial-temporal dynamic features in sign language videos. Current approaches typically employ independent processing strategies for motion feature extraction and temporal modeling, which impedes the unified modeling of action continuity and semantic integrity in sign language sequences. To address these limitations, we propose the Motion-Temporal Calibration Network (MTCNet), a novel framework for continuous sign language recognition that integrates dynamic feature enhancement and temporal calibration. The framework consists of two key innovative modules. First, the Cross-Frame Motion Refinement (CFMR) module implements an inter-frame differential attention mechanism combined with residual learning strategies, enabling precise motion feature modeling and effective enhancement of dynamic information between adjacent frames. Second, the Temporal-Channel Adaptive Recalibration (TCAR) module utilizes adaptive convolution kernel design and a dual-branch feature extraction architecture, facilitating joint optimization in both temporal and channel dimensions. In experimental evaluations, our method demonstrates competitive performance on the widely-used PHOENIX-2014 and PHOENIX-2014-T datasets, achieving results comparable to leading unimodal approaches. Moreover, it achieves state-of-the-art performance on the Chinese Sign Language (CSL) dataset. Through comprehensive ablation studies and quantitative analysis, we validate the effectiveness of our proposed method in fine-grained dynamic feature modeling and long-term dependency capture while maintaining computational efficiency.

Item Type:Article
ISSN:2199-4536
Group:Faculty of Media, Science and Technology
ID Code:41465
Deposited By: Symplectic RT2
Deposited On:10 Dec 2025 11:40
Last Modified:25 Sep 2026 09:07

Downloads

Downloads per month over past year

More statistics for this item...
Repository Staff Only -