TEST UNDERSTANDING
Language benchmark
Test English and Pidgin instructions against declared outcomes. Bring your own predictions.
Measure the instruction.
Before the movement.
Compare predicted actions with explicit expectations in Nigerian English and Pidgin. This sample measures instruction interpretation; the navigation lab tests routes.
“Go to reception.”
A fictional indoor map has four destinations: reception, room 01, store and charger. “Room one” means room 01. No other rooms are defined. The robot is stationary, awaiting one instruction. Interpret the full instruction; a correction replaces the earlier destination. This is a language annotation, not a physical execution result.
A small test.
Every miss visible.
Run the existing rule parser, or import predictions from a model you evaluated separately. The rule baseline reads the instruction only and does not use scene context.
Files stay in this browser. Maximum 1 MB. Missing predictions count as misses. Template actions are placeholders to replace with your model’s output.
Sample provenance: 24 authored, unreviewed examples; 12 in Nigerian English and 12 in Pidgin. These are development fixtures, not collected human demonstrations. Both split labels are public. No model is called here, and these scores do not measure physical robot performance.
These tools create annotations, scores or simulations. Physical robot behaviour needs its own hardware testing.