05 / 05
05
Late 2025 – Jan 2026

Grounded Planning Research

Exposing brittle AI reasoning that high success rates hide.

AI ResearchPlanning SystemsPythonAblation StudyZenodo
View on GitHub Read on Zenodo
The Problem

AI planning benchmarks are saturating. Agents report 90%+ success rates on standard tasks. The assumption is that high success rates mean robust planning, but success rates measure outcomes, not the reasoning that produced them.

An agent can get to the right answer through fundamentally brittle logic, and you won't know until it fails on a slightly different problem. The question I wanted to answer: is there a pressure point that makes the brittleness visible before it fails in deployment?

What I Built

Object commitment is that pressure point. When you force an agent to explicitly name which objects it's operating on before selecting actions (rather than letting it discover them opportunistically) you expose something important: agents that were planning correctly stay correct, and agents that were guessing in the right direction fall apart.

The study runs a controlled ablation in a deterministic filesystem environment. Action-only variant vs. object-centric variant. Same tasks, same success criteria. The failure modes in the object-centric variant are structured and repeatable. You can predict where they'll occur, which means you can actually debug the reasoning.

How it works
01
Environment

Deterministic filesystem tasks: file operations, directory traversal, conditional logic. Clean ground truth, no ambiguity in success/failure.

02
Ablation design

Action-only: agent selects an action and a target is inferred. Object-centric: agent must commit to the specific object before selecting the action.

03
Failure analysis

Object-centric variant surfaces consistent failure patterns tied to specific object relationships. These patterns don't appear in aggregate success metrics.

04
Implication

High-level success rates are insufficient diagnostics. Object commitment as a forced step is a practical technique for surfacing latent planning failures before deployment.

More projects
Impact
3
GitHub stars
Jan 9
Published on Zenodo, 2026
Age 15
Written and published at 15
MIT
Open licensed
vassu-v/action-vs-object-planning
shoryavardhaans2@gmail.com →