
Key technical findings. Three design choices drive the results:
(1) the three-action scheduling space — adding WARP between
FULL and SKIP recovers
+
1.7 pp MOTA vs. pure skipping at
minimal cost; (2) phase-correlation warping is training-free,
O(N log N )
, and matches the accuracy of dense Farnebäck warp
at 5
×
lower overhead (Table 5); and (3) the 803-parameter Pol-
icyMLP outperforms hand-tuned thresholds by 0.6–1.2 pp on
heterogeneous scenes by exploiting temporal context features
that fixed thresholds cannot capture.
Practical impact. The per-sequence results (Table 8) confirm
AFSS matches BL-FULL within 2.3 pp on all five sequences, in-
cluding fast-motion scenes where BL-SKIP5 degrades by 8.5 pp.
The overhead table (Table 9) shows that AFSS adds only 3.1 ms
of fixed overhead — less than 1.6% of the FULL frame bud-
get. This makes the framework deployable on Raspberry Pi 4
(
2 → 12
FPS), Jetson Nano (
8 → 40
FPS), and laptop CPUs
(5 → 32 FPS) without GPU.
Limitations and future directions.
•
Standard benchmarks. Validation on MOT17/MOT20 with
public annotations is the primary planned extension for venue
submission.
•
RL policy. A PPO/SAC agent directly optimising
r
t
=
∆MOTA
t
− λC(a
t
)
could exceed the threshold-policy
teacher, particularly on non-stationary scenes.
•
Multiplicative compression. Combining AFSS with INT8
quantisation (
4×
backbone speedup) projects to
20
–
28×
total
acceleration.
•
Moving cameras. Ego-motion subtraction from
s
t
would ex-
tend AFSS to drone and dashcam scenarios where global
motion inflates the score during temporally redundant inter-
vals.
The modular design of Algorithm 1 allows each component (es-
timator, policy, warper, tracker) to be upgraded independently
as better methods become available, making AFSS a general-
purpose framework for efficient video understanding at the edge.
REFERENCES
[1]
Ultralytics: YOLOv8: A new state-of-the-art model for ob-
ject detection. GitHub (2023).
https://github.com/
ultralytics/ultralytics
[2]
Zhang, Y., et al.: ByteTrack: Multi-object tracking by associating
every detection box. In: ECCV, pp. 1–21 (2022)
[3]
Redmon, J., et al.: You only look once: Unified, real-time object
detection. In: CVPR, pp. 779–788 (2016)
[4]
Bewley, A., et al.: Simple online and realtime tracking. In: ICIP,
pp. 3464–3468 (2016)
[5]
Wojke, N., et al.: Simple online and realtime tracking with a deep
association metric. In: ICIP, pp. 3645–3649 (2017)
[6]
Zhu, X., et al.: Deep feature flow for video recognition. In: CVPR,
pp. 2349–2358 (2017)
[7]
Zhu, X., et al.: Flow-guided feature aggregation for video object
detection. In: ICCV, pp. 408–417 (2017)
[8]
Kag, A., et al.: Video understanding with task-specific temporal
grounding. arXiv:2107.05996 (2021)
[9]
Mullapudi, R.T., et al.: Online model distillation for efficient video
inference. In: ICCV, pp. 3573–3582 (2019)
[10] Jacob, B., et al.: Quantization and training of neural networks for
efficient integer-arithmetic-only inference. In: CVPR, pp. 2704–
2713 (2018)
[11]
Bochkovskiy, A., Wang, C.Y., Liao, H.Y.M.: YOLOv4: Optimal
speed and accuracy of object detection. arXiv:2004.10934 (2020)
[12]
Lin, T.Y., et al.: Focal loss for dense object detection. In: ICCV,
pp. 2980–2988 (2017)
[13]
Farnebäck, G.: Two-frame motion estimation based on polynomial
expansion. In: SCIA, pp. 363–370 (2003)
[14]
Bernardin, K., Stiefelhagen, R.: Evaluating MOT performance:
The CLEAR MOT metrics. EURASIP JIVP 2008, 1–10 (2008)
[15]
Ristani, E., et al.: Performance measures and a data set for multi-
target, multi-camera tracking. In: ECCV, pp. 17–35 (2016)
[16]
Q. Wu, R. Cui, Y. Li, and H. Zhu, “HaltingVT: Adaptive
Token Halting Transformer for Efficient Video Recognition”
arXiv:2401.04975, 2024.
[17]
H. Wang, B. Dedhia, and N. K. Jha, “Zero-TPrune: Zero-Shot
Token Pruning Through Leveraging of the Attention Graph in
Pre-Trained Transformers” Proceedings of CVPR 2024, pp.
16070–16079, 2024.
[18]
D. Nimma and A. Uddagiri, “OPT-STViT: Video Recognition
through Optimized Spatial-Temporal Video Vision Transformers,”
2024.
[19]
H. Ding, C. Guo, J. Sun, X. Jiang, H. Shi, and J. Li, “Motion-
Driven Adaptive Frame Selection Strategy for Video Action Recog-
nition” Journal on Image and Video Processing, vol. 2025, no. 12,
2025.
[20]
L. Shen, G. Gong, T. He, Y. Zhang, P. Liu, S. Zhao, and G. Ding,
“FastVID: Dynamic Density Pruning for Fast Video Large Lan-
guage Models” arXiv:2503.11187, 2025.
[21]
J. Zhang, Y. Yang, R. Tripathi, W. Han, R. Krishna, C. Clark, Y.
J. Lee, and S. Lee, “Unified Spatio-Temporal Token Scoring for
Efficient Video VLMs” arXiv:2603.18004, 2026.
8