Skip to content
← Tags

#Object-Level-Memory (1)

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

DGIST team, accepted to IROS 2026. Robots operating long-term revisit evolving environments; existing systems either overwrite history to keep an up-to-date map or store semantic snapshots without cross-session object identity — the authors call this temporal amnesia: unable to answer "Where has the green chair been across all sessions?" LT-Mem is a three-layer solution: (1) Perception: MASt3R-SLAM multi-session alignment (Sim(3) anchors + inter-session loop closures + joint g2o optimization) plus SAM3 instance segmentation extracting per-object centroid/volume/visual embedding; (2) Reasoning: five evidence scores (E1 spatial proximity, E2 temporal continuity, E3 feature similarity as the primary signal w=0.45, E4 motion consistency, E5 occlusion handling) drive deterministic cross-session re-identification, with a constrained LLM judge only for ambiguous cases (MATCH/NEW-TRACK/HOLD); a structural integrity check triggers a session-wide HOLD on alignment failure; a volatility score V accumulates evidence in a Bayesian manner (Eq.2, one-time LLM prior), selecting among OVERWRITE/HOLD/MULTI-HYPOTHESIS updates; (3) Tri-Memory: Live stores current states, Delta logs timestamped events (APPEAR/DISAPPEAR/MOVE/RE-APPEAR/NONE), Meta accumulates volatility statistics feeding back into the update policy. The paper releases LT-VQA (two indoor envs, 10 objects × 10 sessions each, plus an outdoor parking lot over 10 sessions; 61 state-change events and 80 QA pairs). Results: Event F1 0.910 / QA-Event 0.820 / QA-Freq 0.600, beating Geometric/Text-Batch/VLM-Batch/STAR baselines across the board, with 438K tokens — 1/16 of VLM-Batch (7,114K) and 1/100 of STAR (45,859K); swapping in Qwen2.5-3B still reaches 0.885, showing the gains come from the structured memory architecture rather than LLM capacity. Ablations: removing re-identification collapses F1 to 0.140; removing volatility-awareness drops it to 0.615. Scene-level statistics on the parking lot even answer "when should I go to find a parking spot" (8:00 AM, 0 cars vs 1:03 PM, 34 cars).

Yumin Lee, Hyoseok Ju, Giseop KimAug 19, 2026
SLAMLong-Term-AutonomyScene UnderstandingAug 19, 2026