TALK-Dem: Benchmarking Embodied Task Planning under Dementia-Associated Communication Patterns
What happened
arXiv:2609.38371v1 Announce Type: new Abstract: Existing LLM (the kind of AI system trained on text to produce text)-driven robot task planners rely on a taken-for-granted assumption of an ideal user whose instructions are clear, complete, and task-focused. However, when interacting with real-world users, especially those experiencing cognitive impairments, such as people living with dementia (PLWD), the planners often make mistakes and even pose physical safety risks.
TALK-Dem contains 4,800 instructions and covers five typical communication patterns, including Referential Imprecision, Object Substitution, Empty Speech, Topic Drift, and Intrusion, at three intensity levels. Experiments across six open-weight (a model whose trained parameters anyone may download) (published so anyone may download the trained model) LLMs reveal a substantial robustness gap.
Across communication patterns, open-weight models exhibited performance drops of up to 22.3 percentage points compared to ideal instructions. CARE generally outperformed standard prompting baselines across the six open-weight models, improving average task success by 18.1 percentage points over the vanilla prompt.
Key facts
- TALK-Dem — includes: 4,800 instructions and covers five typical communication patterns, including Referential Imprecision, Object Substitution, Empty Speech, Topic Drift, and Intrusion, at three intensity levels
Sources & evidence
- arXiv Robotics (cs.RO) Reporting source
TALK-Dem: Benchmarking Embodied Task Planning under Dementia-Associated Communication Patterns ↗
https://arxiv.org/abs/2609.38371