NeurIPS 2026

OntoPlan

An Ontology-Grounded Scene Representation and Agentic Framework for Scalable Robot Task Planning

Hyeongwoo Nam Woongje Cho* Juwon Kim* Jongeun Choi†

School of Mechanical Engineering, Yonsei University, Seoul, Republic of Korea * Equal contribution    † Corresponding author

S01Klickitat · goal

I think I left my cell phone in the dining room on Floor B. Can you find it and bring it to the home office on Floor C?

cell_phone_61 · home_office_18

01Overview

From an instruction to an executable plan

  • LLM agents plan over an ontology-grounded world model whose classes and properties map one-to-one to a PDDL domain, so grounding, goals, and planning share one vocabulary.
  • They query only the task-relevant part of the scene and hand formalized goals and constraints to a symbolic planner, so returned plans are executable.
  • On 150 long-horizon tasks in five buildings at three scales: 0.89 task success vs. 0.27 for the strongest baseline, with 5.6× fewer tokens, and cost stays nearly flat as scenes grow 10–23×.
Timeline of an OntoPlan interaction: the user asks to put coffee in the living room, the Scene Explorer finds no coffee but a coffee machine, the Flow Orchestrator asks the user, the Task Formalizer writes conditions, the Planning Manager produces a plan with the PDDL plan tool, and the robot executes it.

The Flow Orchestrator routes grounding, clarification, task formalization, and planning across specialized agents, while the Scene query and PDDL plan tools operate over a shared ontology-grounded world model.

Read the abstract

Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions.

We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans.

Across 150 general tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.89 average task success, compared with 0.27 for the strongest baseline, while using 18.1k total tokens per task on average, about 5.6× fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in success and remain far more costly in tokens. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans.

02Results at a glance

Higher success at a fraction of the token cost

0.89
average task success on 150 general tasks, vs. 0.27 for SayPlan and 0.13 for DELTA
5.6×
fewer total tokens per task than the most efficient baseline (18.1k vs. 100.6k)
1.15×
growth in tokens per task from Small to Large scenes, while object counts grow 10–23×

Task success

Mean over 5 environments × 3 scene scales (higher is better)

SayPlan
0.27
DELTA
0.13
OntoPlan
0.89

Total tokens per task

Input + output tokens over all LLM calls, ×10³ (lower is better)

SayPlan
216.0k
DELTA
100.6k
OntoPlan
18.1k
03Method · World model

One symbolic state for grounding and planning

The world model has three layers: an ontology layer declaring entity types and relation semantics, a static layer storing topology and connectivity, and a dynamic layer storing object instances, states, and spatial relations. Access goes through an ask/tell interface. Ask runs read-only SPARQL queries for scene inspection and planner-state extraction. Tell inserts and deletes RDF triples for observations and action effects, and GraphDB materializes the resulting OWL 2 RL entailments.

World model with ontology, static, and dynamic layers, the ask and tell interfaces, and the OWL reasoner.

World model. Ask supports scene inspection and planner-state extraction; tell updates dynamic facts from symbolic observations and action effects, and GraphDB materializes the OWL 2 RL entailments. The Marstons Medium scene has 2,908 asserted triples and 19,394 after reasoning.

One vocabulary from the world model to the planner

World model · OWL / RDFschema.ttl
:isInsideOf a owl:ObjectProperty ,
                owl:TransitiveProperty ;
    rdfs:domain :Artifact ;
    rdfs:range  :Artifact .

# dynamic layer fact
:toothbrush_213 :isInsideOf :storage_cabinet_55 .
Planner · PDDLdomain.pddl
(:types ... Artifact)

(:predicates
  (isInsideOf ?x - Artifact ?c - Artifact)
  ...)

; extracted initial state
(isInsideOf toothbrush_213 storage_cabinet_55)

Ontology classes map one-to-one to PDDL types, and ontology properties map one-to-one to predicates with the same names. Grounding, task formalization, and planning share one vocabulary, so no schema translation sits between them.

04Method · Agentic workflow

Four roles, two tools, shared memory

OntoPlan splits the planning loop into workflow control, scene grounding, task formalization, and executable planning. Shared memory holds the current intent, grounding evidence, open ambiguities, formalized goals and constraints, and planner diagnostics, so any agent can revise an intermediate result as planning proceeds.

System architecture: the Flow Orchestrator, Scene Explorer, Task Formalizer, and Planning Manager each read and write fields of a shared memory and have their own routing options.
05Demo · Recorded runs

Watch OntoPlan solve a task

Four runs from the evaluation, replayed step by step: the agents' messages as logged, the conditions and subgoals they produced, and the plan executed on the Klickitat scene graph. Nothing here is regenerated; every message and action comes from the run's result file.

The purple marker is the robot and its trail is the executed path; orange squares are objects being carried or put down, and a red cross marks a place an always condition forbids. In the first run the Flow Orchestrator asks which couch to use and receives the fixed automatic reply used in the non-interactive evaluation.

06Results · Scalability

Success holds and token cost stays flat as scenes grow

From Small to Large, object counts grow from 50–121 to 954–1,269 per environment. OntoPlan keeps the highest rate on every metric at every scale, while SayPlan shows heavier token tails and DELTA remains far more token-intensive.

SayPlanDELTAOntoPlan S / M / L = Small / Medium / Large scene
Success rates by scene scale
Overall, 150 tasks
Token usage by scene scale (log scale)
Mean, S / M / L
Per-environment results (table)

Rates are averaged over the three scales; tokens are ×10³. All methods use GPT-4o at temperature 0 and are scored against the same at_end, always, sometime, and ordering checks.

EnvironmentMethodPlan
Returned
Plan
Executable
Goal
Reached
Task
Success
Input tokens
per call
Total tokens
per task
KlickitatSayPlan0.670.470.330.2730.9221.7
DELTA0.630.230.230.2023.9101.3
OntoPlan1.001.000.930.902.617.3
LakevilleSayPlan0.570.400.330.3333.3205.5
DELTA0.470.230.170.1723.498.8
OntoPlan1.001.000.970.972.618.8
LindenwoodSayPlan0.630.470.400.4034.5206.6
DELTA0.370.170.230.1723.498.9
OntoPlan1.001.000.870.872.617.5
MarstonsSayPlan0.730.430.300.2737.1255.6
DELTA0.270.170.170.1023.9101.3
OntoPlan1.001.000.770.772.618.7
MuleshoeSayPlan0.700.170.200.1031.9190.4
DELTA0.470.030.030.0324.2102.7
OntoPlan0.970.970.930.932.618.1

Ablation: what each tool contributes

Without Scene query, the full scene is serialized into the prompt. Without PDDL plan, the LLM generates action sequences directly, given the same PDDL domain. Tokens are ×10³; L/S is the Large-to-Small ratio.

SettingTask successTotal tokens per taskPeak input tokens per call
SMLAvg.SMLL/SSMLL/S
Full system0.920.900.840.8917.117.519.71.15×3.63.73.91.06×
w/o Scene query0.860.900.840.8718.132.572.74.02×7.622.362.28.13×
w/o PDDL plan0.320.260.360.3123.625.026.91.14×8.48.58.41.01×

Selective scene access drives token scalability; the integrated planning path drives task reliability.

07Results · Ambiguous and infeasible instructions

Asking before acting

Two qualitative subsets of 10 instructions each in Klickitat (Medium). When an instruction leaves the target open, OntoPlan asks the user; when a request cannot be satisfied, it reports what is missing instead of committing to a wrong plan.

Ambiguous instructions

MethodPlan ReturnedPlan ExecutableGoal Reached
SayPlan5/104/100/10
DELTA5/103/100/10
OntoPlan (auto replies)8/108/101/10
OntoPlan (user replies)10/1010/108/10

Auto replies: a fixed answer telling the agent to assume and continue, matching the non-interactive baselines.

Infeasible instructions

MethodClarification
request
Misdirected
plan
Inexecutable
plan
Planning
failure
SayPlan—3/100/107/10
DELTA—1/101/108/10
OntoPlan9/101/100/100/10

SayPlan and DELTA cannot request clarification in this protocol.

See the exchange for each instruction

    08Qualitative · One session, three turns

    A multi-turn session over one world model

    A question, a task, and a follow-up in one session. Expand a tool call to see its arguments and result. Plans run on the shared world model, and each turn ends with the facts it changed, so the next turn starts from the updated state. Recorded with the released code; the paper's figure shows an earlier run of the same session, so details such as where the mug is put down differ.

    Dashed blue lines are paths returned by find_path; the purple trail is the executed plan. In Turn 3 the robot starts in the living room where Turn 2 left it, and the bread is still in the switched-on oven.

    09Benchmark

    Five buildings, three scene scales

    Environments come from the 3D Scene Graph dataset and are mapped deterministically into one ontology schema (Small). Medium is manually augmented, and Large triples Medium's object count through constrained random augmentation. All scales share the same ontology and PDDL domain, and all are included in the repository.

    EnvironmentFloorsRoomsSmallMediumLarge
    Klickitat328843851,155
    Lakeville219833871,161
    Lindenwood23659318954
    Marstons431503881,164
    Muleshoe2241214231,269

    Small / Medium / Large columns are object counts.

    10Citation

    BibTeX

    @inproceedings{nam2026ontoplan,
      title     = {{OntoPlan}: An Ontology-Grounded Scene Representation and
                   Agentic Framework for Scalable Robot Task Planning},
      author    = {Nam, Hyeongwoo and Cho, Woongje and Kim, Juwon and Choi, Jongeun},
      booktitle = {Advances in Neural Information Processing Systems},
      year      = {2026}
    }