APM-Bench Benchmarking cross-Session Persistent Memory for
Real-World Egocentric Streaming Video Assistants

Jianguo Huang1,2,* Jinming Liu1,2,* Qiyao Wang3 Liang Xu4 Jianhang Li5 Zhimian Wen2
Mingda Li5 Shule Lu6 Zhicheng Wang2,7 Yuhan Guo1,2 Xin Jin2 Wenjun Zeng2,†
1Shanghai Jiao Tong University 2Eastern Institute of Technology, Ningbo
3Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
4Zhongguancun Academy, Beijing, China 5Dalian University of Technology
6Beihang University 7Hong Kong Polytechnic University

* Equal contribution.† Corresponding author.

Your browser cannot display this PDF inline.

In a streaming setting, once an interaction ends, the model can no longer directly access its visual stream because real-world interactions are not replayed; the assistant therefore needs persistent memory to retain prior experience. Across sessions, the assistant updates and reuses persistent memory for cross-session understanding and proactive assistance while continuing real-time perception. The example shows repeated collaborative dessert-making across multiple sessions.

Abstract

To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended questions. Each session is a video with fine-grained annotations, and sessions within a trajectory revolve around related activities. Models then use persistent memory to answer questions about past sessions and provide proactive responses while maintaining real-time interaction. This raises challenges: persistent memory must be storable, selectively retain information, be injected at the right time, and remain efficient. Moreover, finite storage may leave required evidence unavailable, so assistants should recognize missing evidence. Therefore, we systematically evaluate general video models under different memory protocols and diverse specialized streaming memory systems, and test whether models acknowledge insufficient evidence. Our evaluation reveals a clear utility–latency–storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions. APM-Bench provides a comprehensive testbed for developing and comparing persistent memory systems under realistic streaming conditions. We hope it encourages future work that jointly considers utility, latency, and storage toward more practical persistent memory for real-world streaming assistants.