An assistant can produce fluent Pidgin while choosing the wrong action. Task evaluation asks a narrower question: given this instruction and this scene, did the system identify what should happen? That distinction matters when a language response will eventually influence navigation or physical work.
Specify the scene and the answer
For an indoor navigation example, name the available destinations and explain relevant aliases. Then label the expected action: navigate to a known destination, stop, or request clarification. A correction may replace an earlier destination. A request for an undefined place may require clarification rather than a confident guess.
Evaluate the complete instruction together with its context. A keyword match on “reception” can be misleading if the speaker explicitly rules reception out. Likewise, a negation followed by a clear alternative can still define an unambiguous task.
Know what the sample contains
The indoor-language-v0.1 sample contains 24 authored fixtures: 12 Nigerian English examples tagged en-NG and 12 Pidgin examples tagged pcm. Each language includes navigation, stop and clarification cases. These are written examples in a fictional setting, not collected recordings or evidence of population-wide language performance.
The wording and expected labels are marked unreviewed. The public development and evaluation labels help organise diagnostics, but a public evaluation split is not a private, independent test set. Read the dataset notes and rights statement before reusing the sample.
Compare outputs consistently
- Record the sample version and the task identifiers being evaluated.
- Provide each system with the same permitted instruction and context.
- Map its answer into navigate, stop or clarify, with a destination only for navigation.
- Count omitted answers as misses and inspect errors by language and action.
- Keep predictions and the exported scoring report alongside your configuration.
The 9jaRobotics task benchmark runs a small rule baseline and accepts predictions generated elsewhere. The baseline reads the instruction alone and ignores scene context. Its limitations are useful to inspect; its score is not a claim about a trained language model.
Improve the next collection
Use mistakes to identify missing distinctions, then ask qualified speakers to review new examples. Preserve natural phrasing, annotate the intended action and document ambiguous cases instead of forcing agreement. Avoid making both your examples and labels imitate the baseline grammar.
When you move to recorded demonstrations, use the task studio to preserve time ranges, scene details and permission metadata. Text interpretation is only one layer: speech recognition, visual perception and successful robot execution require separate inputs and separate measurements.
Bring people into the project
Use 9jaTesters to request collection, annotation, testing or human review. Describe the task, volume, location or languages, and acceptance criteria in your brief. For contributor opportunities, join the 9jaTesters workforce.