Hu, H., Peng, J., Xiao, Z., Guo, L. and Wu, D., 2025. Motion-temporal calibration network for continuous sign language recognition. Complex and Intelligent Systems, 12, 35 (2026).
Full text available as:
Preview |
PDF (OPEN ACCESS)
s40747-025-02156-5.pdf - Published Version Available under License Creative Commons Attribution. 2MB |
|
PDF
Title Page.pdf - Accepted Version Restricted to Repository staff only 151kB | |
|
PDF
MTCNet.pdf - Accepted Version Restricted to Repository staff only 1MB | |
|
Copyright to original material in this document is with the original owner(s). Access to this content through BURO is granted on condition that you use it only for research, scholarly or other non-commercial purposes. If you wish to use it for any other purposes, you must contact BU via BURO@bournemouth.ac.uk. Any third party copyright material in this document remains the property of its respective owner(s). BU grants no licence for further use of that third party material. |
DOI: 10.1007/s40747-025-02156-5
Abstract
Continuous Sign Language Recognition (CSLR) is fundamental to bridging the communication gap between hearing-impaired individuals and the broader society. The primary challenge lies in effectively modeling the complex spatial-temporal dynamic features in sign language videos. Current approaches typically employ independent processing strategies for motion feature extraction and temporal modeling, which impedes the unified modeling of action continuity and semantic integrity in sign language sequences. To address these limitations, we propose the Motion-Temporal Calibration Network (MTCNet), a novel framework for continuous sign language recognition that integrates dynamic feature enhancement and temporal calibration. The framework consists of two key innovative modules. First, the Cross-Frame Motion Refinement (CFMR) module implements an inter-frame differential attention mechanism combined with residual learning strategies, enabling precise motion feature modeling and effective enhancement of dynamic information between adjacent frames. Second, the Temporal-Channel Adaptive Recalibration (TCAR) module utilizes adaptive convolution kernel design and a dual-branch feature extraction architecture, facilitating joint optimization in both temporal and channel dimensions. In experimental evaluations, our method demonstrates competitive performance on the widely-used PHOENIX-2014 and PHOENIX-2014-T datasets, achieving results comparable to leading unimodal approaches. Moreover, it achieves state-of-the-art performance on the Chinese Sign Language (CSL) dataset. Through comprehensive ablation studies and quantitative analysis, we validate the effectiveness of our proposed method in fine-grained dynamic feature modeling and long-term dependency capture while maintaining computational efficiency.
| Item Type: | Article |
|---|---|
| ISSN: | 2199-4536 |
| Group: | Faculty of Media, Science and Technology |
| ID Code: | 41465 |
| Deposited By: | Symplectic RT2 |
| Deposited On: | 10 Dec 2025 11:40 |
| Last Modified: | 25 Sep 2026 09:07 |
Downloads
Downloads per month over past year
| Repository Staff Only - |
Tools
Tools