Skip to main content

Do transformer temporal heads and post-pooling motion gates help CorrNet-based CSLR? An empirical study.

Wang, L., Xiao, Z. and Peng, J., 2026. Do transformer temporal heads and post-pooling motion gates help CorrNet-based CSLR? An empirical study. ArXiv.

Full text available as:

[thumbnail of Do Transformer Temporal Heads and Post-Pooling Motion Gates Help CorrNet-based CSLR.pdf]
Preview
PDF
Do Transformer Temporal Heads and Post-Pooling Motion Gates Help CorrNet-based CSLR.pdf - Accepted Version

315kB

Official URL: https://doi.org/10.48550/arXiv.2607.09890

DOI: 10.48550/ARXIV.2607.09890

Abstract

CorrNet is a strong baseline for continuous sign language recognition (CSLR) because it models inter-frame correlations inside the visual encoding stage. In this paper, we study two natural extensions of a reproduced CorrNet system: replacing the BiLSTM temporal head with a Transformer encoder, and injecting motion cues after temporal pooling. We find that the Transformer head does not outperform the BiLSTM baseline, even with a training strategy adjusted for the Transformer, and the two heads have almost the same computational and runtime cost. For the second extension, we design a lightweight module called MotionGate. In our experiments, MotionGate consistently collapses to an identity-like mapping: the gate loses motion selectivity, and the injected residual becomes a weak, non-selective perturbation of the pooled features. These results suggest that explicit motion injection after CorrNet’s correlation-based encoding is largely redundant, and that natural-looking architectural extensions in CSLR should be tested carefully instead of being assumed to help.

Item Type:Other
Group:Faculty of Media, Science and Technology
ID Code:42322
Deposited By: Symplectic RT2
Deposited On:24 Aug 2026 13:21
Last Modified:24 Aug 2026 13:21

Downloads

Downloads per month over past year

More statistics for this item...
Repository Staff Only -