Abstract:Referring Video Object Segmentation (R-VOS) segments specified targets in videos according to language descriptions. However, its performance in complex dynamic scenes is hampered by uneven frame quality and the temporal instability of DETR-based methods. To tackle these challenges, a Multi-Scale Quality-Aware Tracking method (MS-QAT) is proposed. Based on ReferDINO, MS-QAT integrates two plug-and-play modules: a Multi-Scale Temporal Quality-Weighted Fusion module (MST-QWF) to suppress low-quality frame interference by dynamic weighting, and a Multi-Scale Temporal Tracking Smoothing Fusion module (MST-TSF) to mitigate tracking drift through temporal smoothing. MS-QAT boosts the mAP by 1.9% on A2D-Sentences and 0.9% on JHMDB-Sentences compared to the baseline. Ablation studies further confirm that MS-QAT effectively enhances segmentation accuracy, temporal stability, and overall robustness across different baselines, offering a potent solution for semantically guided object segmentation in dynamic videos.