Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
Abstract
Recent advances in robot learning have enabled manipulation policies to perform increasingly diverse tasks and generalize across environments. However, reliable execution often depends on hidden task states that cannot be determined from current observations alone, making interaction history essential. We introduce HIDE, a benchmark for evaluating manipulation memory under partial observability. HIDE comprises 15 tasks covering repetition counting, historical-state recall, and execution-progress tracking, with randomized initial configurations and decision points where similar observations require different actions depending on prior events. We further propose SEEK, a framework combining three complementary memory mechanisms to retain historical evidence and track execution state. Evaluations reveal substantial limitations in existing policies on HIDE, while memory augmentation improves task success in both simulation and real-world experiments. Individual mechanisms benefit some tasks but can degrade others; their combination achieves the highest average success rate on HIDE among the evaluated configurations. These findings highlight the importance of maintaining internal representations of hidden task states and matching memory design to task-specific information requirements.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation (2026)
- Simple Agentic Memory for Generalist Robot Policies (2026)
- MEMOBench: A Process Level Memory Benchmark for Robotic Manipulation (2026)
- MemBodied: Recurrent Associative Memory for Vision-Language-Action Models (2026)
- Memory as Plans: World-Action Modeling with Memory-Grounded Planning (2026)
- ActiveArena: Benchmarking and Understanding Active Perception in Robotic Manipulation (2026)
- HarnessWAM: Bridging Prediction and Deliberation in World Action Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.38886 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper
