BEYOND MEMORY / LONG-HORIZON AGENTS

Progression
of States.

Beyond memory. Toward progress.

Remembering what happened is not enough. An agent must understand what has changed, recognize when progress stalls, and recover.

ARXIV PREPRINT / OCTOBER 2026

Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States

Yu Luo1, Jiamin Jiang1, Yimin Zuo1, Xidao Wen2, Rongchen Gao1, Yongqian Sun1,*, Shenglin Zhang1, Guiyang Liu2, Cheng Zhang2, Fang Situ2, Qi Zhou2, Dan Pei3

1Nankai University 2Alibaba Group 3Tsinghua University

* Corresponding author: Yongqian Sun

PoS performance across four benchmarks and an ALFWorld case: Raw Trajectory fails after 50 steps; explicit belief and recovery succeed in 17 steps.
A long interaction history is not the same as progress. PoS keeps the current state and the remaining requirements explicit.

01 BEYOND RETAINING HISTORY

Abstract.

Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.

FROM STRUCTURED SEARCH TO MEASURED PROGRESSIONGraph of States grounds abductive search in a causal graph and a state machine. PoS maintains a task-conditioned belief for both diagnosis and execution, assesses how that belief evolves, and intervenes when progress stalls.

02 STATE PERCEPTION

Beyond memory.
Understand what holds now.

Memory retains past interactions. A validated belief describes the current world and what remains unresolved. Progression compares these states over time, so repeated activity can be recognized and redirected.

Beyond Memory: past interactions form a validated current belief; repeated inspection leaves the mug cold and its active gap unresolved; progression monitoring detects stagnation and redirects the next action toward heating.
The same observation can be recorded accurately many times while the task remains unfinished. Full-resolution figure

PoS connects the current belief to the next decision through a repeated update, check, and recovery loop. See the complete method

03 THE COMPLETE LOOP

Inside PoS.

PoS maintains a current belief, checks its consistency, and measures its evolution across time. The Task Agent selects actions using that belief, an active gap, and recovery constraints when they are needed.

ObserveUpdate & validateMeasure progressRecover when trapped
PoS framework: belief modeling, belief-guided interaction, the Belief Sentinel, and trapping-aware recovery
Figure 2. Candidate updates are checked before they guide the next action or contribute to progress monitoring.
WORLD STATE

What holds now

Entities, their states, and their relations, with provenance, confidence, and evidence when available.

EPISTEMIC GAP

What needs to be learned

A question that remains unresolved, such as distinguishing database waiting from JVM-side processing.

ACHIEVEMENT GAP

What needs to change

A difference between the current world and the goal, such as making a mug hot before putting it away.

CONSISTENCY COMES FIRST

Before measuring progress, check the belief.

A candidate update can contain conflicting states or claims unsupported by interaction evidence. The Belief Sentinel returns issues for revision; only the validated belief becomes the next decision context.

Two incompatible states in one update.

The candidate records the microwave as both open and closed. It is not ready to guide the next action.

Schematic update illustrating Section 3.2, not a recorded case step.

04 PROGRESSION ACROSS TIME

Activity is visible.
Progress must be assessed.

PoS checks state updates against evidence before measuring them. It examines what remains unresolved, whether useful changes occur, and whether earlier states recur. Action counts alone cannot tell us when the task needs a different direction.

PROGRESS MONITOR

Two windows of the same task
ACTIVE GAPMake the mug hot.One of two initial requirements remains
Eight validated transitionsWindow ready
Recorded progressNo recorded progressLetters describe the state relevant to the current requirement.
H=1−p⋅max(S,R)

1 − 0.50 × max(0.50, 0.57) = 0.71

Stronger trapping signalHealthier belief dynamics
WHY COMBINE THESE SIGNALS?

A requirement can remain open while progress continues.

Gap persistence alone does not establish trapping. PoS combines it with stagnation or recurrence.

Inspect one transition
Before

The mug is cold and inside the microwave.

After

The mug remains cold.

The next section uses this task and active gap.Analyze this window
Follow the calculation

Single-step progress

dtdiag=12∑x|ct(x)−ct−1(x)|

Diagnosis records progress when this change exceeds 0.30. Confidence values are maintained independently; decreases can contribute too. In execution, the Sentinel assesses whether a transition reduces the gap, acquires needed information, or advances a plausible path.

Persistence and stagnation

PX=|∩τ∈WΔτX||Δt−K+1X|St=1−1K∑τ∈Wuτ

Compute persistence separately for epistemic and achievement gaps; an empty initial set gives zero. A gap that disappears and later returns is not persistent throughout the window.

Gap-relative recurrence

dW=13∑Y∈{E,S,R}dJ(YiΔ,YjΔ)

Project historical world states onto the same current active gap. Compare aligned Entity, State, and Relation sets over lags 1–4; R is the highest repeat rate. The distance threshold is 0.15. Real belief comparison includes confidence changes.

Section 3.3, Appendix B.2, and Table 8. Detection starts after K = 8 validated transitions and triggers at H ≤ 0.25. The worked examples use aligned symbolic records and the paper's thresholds.

Why this health formula?

Persistence must coincide with stagnation or recurrence. A maximum over all three signals can mistake unfinished work for trapping; a geometric mean can miss trapping when one signal is absent.

Appendix B.3 / Table 4: boundary configurations
Situation(p, S, R)Geometric meanMaximumPoS
Stagnation without recurrence(1, 1, 0)100
Recurrence despite positive step labels(1, 0, 1)100
Progress before gap closure(1, 0, 0)101

These are illustrative boundary cases, not episode measurements. Lower H indicates stronger trapping signals.

05 TRAPPING-AWARE RECOVERY

Recognize the trap.
Restore progress.

Recovery combines two questions: how is the agent stuck, and what remains unresolved? The resulting constraints redirect action selection without supplying the answer or prescribing a tool call.

FACTORIZE. COMPOSE. REASSESS.

Pattern × unresolved requirement

The same task and validated window from Progression.

HOW IS THE AGENT STUCK?
WHAT PROGRESS IS BLOCKED?
ACTIVE GAP / FIXED DURING RECOVERYMake the mug hot.

The relevant world stays unchanged while actions continue.

PATTERN CONSTRAINT

Avoid repeating the ineffective action under the unchanged belief.

GAP CONSTRAINT

Acquire evidence that distinguishes the remaining explanations.

Ct=Cpattern∪CgapConstraints awaiting composition
  1. 01Detect
  2. 02Diagnose
  3. 03Compose
  4. 04Act
  5. 05Validate
  6. 06Reassess
After the next validated observation
TRAPPING DETECTED

A persistent gap meets stalled belief dynamics.

Once a full window is available, H at or below 0.25 triggers trapping diagnosis.

H ≤ θH

Recompute health after each validated transition. Keep or update the constraints while trapping persists.

Mechanism illustration of Section 3.3. Static, Cycle, and Drift are tested in order; unmatched cases use generic recovery. Constraints guide action selection while the active gap remains fixed. The actual diagnostic recovery unfolds in the case below.

06 SEE THE DIFFERENCE UNFOLD

Case Walkthrough.

Follow the same task through two decision contexts. In PoS, inspect the current requirement, the evidence that changes the belief, and the recovery that restores progress. Raw Trajectory retains the interaction record.

ALFWorld #0096

Put a hot mug in the cabinet.

Reported outcome: 17 actions · task completed

Task

A goal with two requirements.

01 / 16
Frame 1 of 16Display frame · mechanism view
Choose a display frame
Evidence and decision detail
What happened

Put a hot mug in the cabinet.

Why it matters

Temperature and location are separate requirements; completing one does not complete the task.

Inspect this step

Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.

Case source and action indices

Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.

Read all four trajectories as text

Execution / Raw Trajectory

Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.

  1. One goal, written into the history. Task

    The task asks for a hot mug in the cabinet.

    Both temperature and location matter, even though the instruction is a single sentence.

    Action: Receive task
    Observation: Put a hot mug in the cabinet.

  2. The mug enters the microwave. Action

    The baseline puts the mug into the microwave.

    Placing the mug in an appliance does not itself heat it.

    Action: Put mug in microwave
    Observation: The mug is in the microwave.

  3. The heating action is missed. Omission

    The baseline moves on without completing the heating requirement.

    The action history contains placement, but no completed heating action.

    Action: Continue without heating
    Observation: The heating requirement has not been fulfilled.

  4. The mug is moved to the cabinet. Action

    The unheated mug is placed in the target cabinet.

    The location now looks right, but the task is still incomplete.

    Action: Put mug in cabinet
    Observation: The mug reaches the cabinet without being heated.

  5. Examine the cabinet. Observation

    The baseline examines the cabinet again.

    Another look cannot cause the missing temperature change.

    Action: Examine cabinet
    Observation: The mug remains in the cabinet.

  6. The next interaction adds little. Repetition

    The same cabinet-examination pattern continues.

    The context grows while the task-relevant state stays unchanged.

    Action: Examine cabinet again
    Observation: No heating action is added to the trajectory.

  7. Thirty-six cabinet checks. Stalled

    The paper summarizes 36 cabinet examinations in this run.

    These repeated observations are activity, not progress toward a hot mug.

    Action: Repeat cabinet examination
    Observation: 36 cabinet checks in the reported trajectory.

  8. The run ends without satisfying the goal. Outcome

    The baseline fails after 50 steps.

    The mug reached the right place, but never reached the required state.

    Action: End of run
    Observation: Failed after 50 steps.

Execution / PoS

Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.

  1. A goal with two requirements. Task

    Put a hot mug in the cabinet.

    Temperature and location are separate requirements; completing one does not complete the task.

    Current understanding: The task specifies temperature and location requirements.
    Unresolved: Make Mug 3 hot, then put it in Cabinet 1.
    Guidance: Act on the current state and the remaining requirements.

  2. Ground the goal in the world. World model

    Mug 3, Microwave 1, and Cabinet 1 become task-relevant entities.

    An entity–state–relation view makes the objects and their current relationships explicit.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Make Mug 3 hot, then put it in Cabinet 1.
    Guidance: Act on the current state and the remaining requirements.

  3. Keep current state separate from history. Belief

    The belief records what is currently known about the mug and the appliances.

    Past actions are evidence for a current state, not a substitute for that state.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Make Mug 3 hot, then put it in Cabinet 1.
    Guidance: Act on the current state and the remaining requirements.

  4. Expose both unmet requirements. Gaps

    The mug needs to become hot and end up in the cabinet.

    PoS carries unresolved requirements alongside the world model.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: 1. Make Mug 3 hot. 2. Place the hot mug in Cabinet 1.
    Guidance: Act on the current state and the remaining requirements.

  5. Choose the heating gap first. Active gap

    The temperature requirement becomes the active gap.

    Location alone would not establish the full goal.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Active: make Mug 3 hot. Pending: place it in Cabinet 1.
    Guidance: Act on the current state and the remaining requirements.

  6. Interaction continues. Heating does not. Low progress

    Further activity leaves the heating requirement unfulfilled.

    The display expands the reported stagnant state; it does not invent an exact early tool sequence.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Active: make Mug 3 hot.
    Guidance: Act on the current state and the remaining requirements.

  7. A consistent belief can still be stuck. Validate

    The Sentinel reports no consistency issue in the illustrated state.

    Validation checks reliability. Progress monitoring asks a different question: is the task moving forward?

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Active: make Mug 3 hot.
    Guidance: Act on the current state and the remaining requirements.

  8. Monitor validated transitions. Monitor

    The active requirement persists across low-progress interactions.

    PoS looks at gap persistence, progress stagnation, and recurrence over a recent window.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Active: make Mug 3 hot.
    Guidance: Act on the current state and the remaining requirements.

  9. Recognize Static + Achievement Gap. Stall

    The reported belief health is 0.1; the missing temperature change remains.

    The system diagnoses both how the task is trapped and what kind of requirement remains.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Active: make Mug 3 hot.
    Guidance: Act on the current state and the remaining requirements.

  10. Combine two recovery constraints. Recovery

    Break the ineffective transition and cause a goal-relevant state change.

    The constraints guide the Task Agent. They do not replace its choice of action.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Active: make Mug 3 hot.
    Guidance: Break the unchanged-state pattern + cause the missing heating transition.

  11. Select a route that can change temperature. Choose action

    The recovery route uses the microwave to heat the mug.

    The active gap stays fixed until supporting evidence establishes that it has been resolved.

    Current understanding: Mug 3 is not yet hot and at the counter.
    Unresolved: Active: make Mug 3 hot.
    Guidance: Address the heating requirement before final placement.

  12. Place the mug in the microwave. Action

    Mug 3 enters Microwave 1.

    The world model updates its location; the heating gap is not yet resolved.

    Current understanding: Mug 3 is not yet hot and in Microwave 1.
    Unresolved: Active: make Mug 3 hot.
    Guidance: A location change alone is not evidence of heating.

  13. Carry out the heating action. Action

    The selected action makes the mug hot.

    A goal-relevant world-state change now supplies a candidate belief update.

    Current understanding: Mug 3 is being heated and in Microwave 1.
    Unresolved: Active: establish that Mug 3 is hot.
    Guidance: Act on the current state and the remaining requirements.

  14. Validate heating, then advance the gap. Validate

    The current belief records the hot mug.

    Only after the heating requirement is resolved does placement become the next active requirement.

    Current understanding: Mug 3 is hot and in Microwave 1.
    Unresolved: Active: put the hot mug in Cabinet 1.
    Guidance: Act on the current state and the remaining requirements.

  15. Move the hot mug to the target. Action

    The heated mug is placed in Cabinet 1.

    The relation changes while the completed temperature requirement is retained.

    Current understanding: Mug 3 is hot and in Cabinet 1.
    Unresolved: Check the final hot-mug-in-cabinet state.
    Guidance: Act on the current state and the remaining requirements.

  16. Both requirements are satisfied. Outcome

    PoS completes the reported task in 17 actions.

    The final belief establishes temperature and location together, rather than mistaking repeated activity for completion.

    Current understanding: Mug 3 is hot and in Cabinet 1.
    Unresolved: Resolved: Mug 3 is hot and in Cabinet 1.
    Guidance: Act on the current state and the remaining requirements.

Diagnosis / Raw Trajectory

Illustrative Raw Trajectory contrast based on Appendix F's diagnostic question. The paper provides no matched baseline trace for t024. These are explanatory interaction summaries, not tool logs or scored results.

  1. Begin with the diagnostic request. Task

    This contrast uses the same inventory-service diagnostic question.

    The following history illustrates Raw Trajectory; it is not a matched baseline run reported in the paper.

    Action: Receive diagnostic request
    Observation: Investigate the inventory-service delay.

  2. Read the initial symptoms. Observe

    Slow requests and elevated GC activity enter the interaction history.

    These clues alone do not settle which explanation is responsible.

    Action: Inspect incident symptoms
    Observation: Slow inventory requests coincide with elevated GC activity.

  3. Consider two possible explanations. Reason

    The agent may consider database waiting and JVM-side processing in its reasoning.

    Reasoning in the transcript is not an explicitly maintained belief-state object.

    Action: Consider explanations
    Observation: Database waiting or processing inside the inventory JVM?

  4. Search for more log evidence. Investigate

    A log search adds another interaction to the context.

    A longer transcript is useful only when the new evidence resolves something.

    Action: Search logs
    Observation: The new material does not distinguish the two explanations.

  5. Continue the same line of investigation. Investigate

    The illustration adds a further search with little new diagnostic value.

    Earlier clues remain available, but no structured gap is maintained beside the history.

    Action: Continue log investigation
    Observation: The diagnostic distinction remains open.

  6. Append another low-information result. Repeat

    Additional interaction can repeat facts that are already in the context.

    The history records the activity without making it a new piece of discriminating evidence.

    Action: Inspect further log material
    Observation: No additional distinction is established in this schematic.

  7. Revisit earlier observations. Review

    The same GC clue can be read again without resolving the cause.

    Retaining a clue is different from obtaining evidence that separates explanations.

    Action: Review earlier observations
    Observation: GC remains a clue; the distinction is still open.

  8. Another search enters the record. Investigate

    A further search adds another interaction.

    This contrast illustrates context growth, not a measured baseline outcome.

    Action: Search for further log information
    Observation: No discriminating evidence is established in this illustration.

  9. More history. The same question. Stalled

    Both explanations are still plausible in this illustrative sequence.

    Raw history retention alone does not provide PoS's explicit progress checks and factorized recovery.

    Action: Review accumulated interactions
    Observation: The database-versus-JVM question is still unanswered.

  10. End of the illustration, not a scored result. Open question

    The schematic ends with the diagnostic question unresolved.

    No baseline failure rate, final diagnosis, or action count is claimed without its actual trace.

    Action: Illustration ends
    Observation: No matched baseline outcome is reported here.

Diagnosis / PoS

Appendix F, p. 26: RCA-100 t024 / Kimi-K3. One recovery episode within a 28-action run. Action indices are zero-based; display frames are not individual tool calls.

  1. Start with the inventory incident. Context

    Inventory-service requests are slow.

    The replay focuses on a recovery episode, not all 28 actions of the reported run.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  2. Carry earlier GC evidence forward. Evidence

    Earlier observations include elevated garbage-collection activity.

    The current belief retains useful evidence without replaying the entire interaction history.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  3. Keep two explanations open. Hypotheses

    Database waiting and JVM-side processing remain competing explanations.

    GC is a clue, not sufficient evidence to decide between every possible cause.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  4. Make the missing distinction explicit. Active gap

    The next investigation needs evidence that separates the alternatives.

    This epistemic gap gives the investigation a specific target.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  5. Search the logs. Investigate

    Further log searches do not settle the distinction.

    More text in the context is not necessarily more diagnostic progress.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  6. Check updates before counting progress. Validate

    Progress monitoring uses validated belief transitions.

    This panel explains the mechanism; no rejected update is claimed for this case.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  7. Notice the persistent gap. Monitor

    Relevant understanding changes little while interactions continue.

    PoS can distinguish an unresolved question from merely a long conversation.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  8. Action 22: diagnose the trap. Stall

    PoS identifies Static Stagnation + Epistemic Gap.

    The diagnosis combines an unproductive transition pattern with the kind of missing requirement.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  9. Break the ineffective pattern. Recovery / pattern

    Avoid continuing the same ineffective investigation.

    The pattern constraint alone does not specify which evidence would resolve the question.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Avoid repeating the ineffective investigation under the unchanged belief.

  10. Add the gap-specific constraint. Recovery / gap

    Seek evidence that discriminates database waiting from JVM-side processing.

    Together, the constraints redirect the agent without handing it a diagnosis or a tool call.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Break ineffective repetition + acquire evidence that separates the alternatives.

  11. The Task Agent chooses CPU inspection. Choose action

    The agent selects CPU investigation.

    This is an agent-selected response to the constraints, not a tool prescribed by the recovery module.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  12. One more low-progress action. Continue

    Recovery initially yields another low-progress interaction.

    A useful constraint need not cause an immediate state change.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  13. The question is still unresolved. Continue

    A second subsequent action leaves the targeted distinction open.

    The active gap is preserved rather than replaced just to make the reasoning appear productive.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  14. Persist through the third action. Continue

    A third low-progress action follows the recovery trigger.

    The replay makes the reported delay visible; it does not invent the individual tool outputs.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  15. New CPU evidence enters the belief. Evidence

    Subsequent CPU queries show increased activity.

    New evidence needs to be interpreted together with the earlier GC observations.

    Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  16. Combine CPU with the earlier GC clue. Revise

    Together, CPU and GC evidence strengthen JVM-side processing over database waiting.

    CPU activity alone does not uniquely establish memoryPressure; the revision belongs to the accumulated evidence.

    Current understanding: CPU + GC observations strengthen the JVM-side explanation.
    Unresolved: Distinguish database waiting from JVM-side processing.
    Guidance: Seek evidence relevant to the active diagnostic question.

  17. Action 26: the targeted gap closes. Progress

    The paper records the epistemic gap as resolved at action 26.

    A supported belief change restores progress toward the final report.

    Current understanding: CPU and earlier GC evidence support JVM-side processing; the targeted gap is resolved.
    Unresolved: Resolved: the targeted database-versus-JVM distinction.
    Guidance: Seek evidence relevant to the active diagnostic question.

  18. Action 27: submit the diagnosis. Outcome

    The agent submits inventory / memoryPressure; both match the benchmark label.

    The final result comes from the full run's evidence, not an isolated observation.

    Current understanding: Final entity: inventory. Final failure type: memoryPressure.
    Unresolved: Resolved: entity and failure type match the benchmark label.
    Guidance: Seek evidence relevant to the active diagnostic question.

Original motivation figure
Figure 1: overall performance and the two independent ALFWorld mug-task trajectories
Figure 1. Raw Trajectory and PoS are separate runs on the same task.

07 TEST THE IDEA

Results.

OVERALL PERFORMANCE

Four benchmarks. Three backbones.

PoS has the highest overall metric among the tested methods in all 12 benchmark-backbone settings. Each comparison below uses the strongest evaluated baseline with the same backbone.

ALFWorld

88.81%

vs. 72.39%

LongHorizon-Harness

LOCA-Bench

56.38%

vs. 52.57%

LongHorizon-Harness

RCA-100

38.83%

vs. 28.16%

PACE / HiAgent

ClinDiag

45.03%

vs. 41.89%

PACE
All evaluated methods and ablations
Table 1 · Qwen3.7-Plus · All values are percentages
MethodALFWorldLOCA-BenchRCA-100ClinDiag
Raw Trajectory62.6943.6224.2738.91
ACON66.4249.927.1839.74
PACE67.9117.3328.1641.89
HiAgent66.4224.5728.1640.07
LongHorizon-Harness72.3952.5726.2140.56
PoS88.8156.3838.8345.03
Without consistency validation73.8844.5731.0744.54
Without trapping diagnosis73.1348.1933.9842.38

Table 1. ALFWorld and LOCA-Bench: overall task success. RCA-100: joint entity-and-type accuracy. ClinDiag: overall diagnosis accuracy. All LLM components use the evaluated backbone; methods share benchmark interfaces and environment-action budgets, not identical trajectories. Baseline adaptations are documented in Appendix D.2.

BEYOND BELIEF CONSTRUCTION

Validation and trapping recovery both matter.

Removing either consistency validation or trapping diagnosis reduces the overall metric in every tested benchmark-backbone setting. The comparison below keeps the backbone fixed.

Backbone: Qwen3.7-Plus

ALFWorld

PoS88.81%

Without validation73.88%

Without trap diagnosis73.13%

LOCA-Bench

PoS56.38%

Without validation44.57%

Without trap diagnosis48.19%

RCA-100

PoS38.83%

Without validation31.07%

Without trap diagnosis33.98%

ClinDiag

PoS45.03%

Without validation44.54%

Without trap diagnosis42.38%

Table 1. Overall benchmark metrics (%); the full method comparison appears above.

Does targeted recovery improve outcomes?

With Qwen3.7-Plus, factorized recovery improves final task performance over a generic recovery prompt on all four benchmarks.

Generic recoveryPoS recoveryFinal task performance (%)
ALFWorldTask success
76.87%
88.81%
LOCA-BenchTask success
49.33%
56.38%
RCA-100Joint accuracy
31.07%
38.83%
ClinDiagDiagnosis accuracy
42.22%
45.03%

Figure 3 (right). These are final outcomes over all evaluated cases, not trap-recovery success rates.

Original trapping and recovery analysis
Figure 3: trapping incidence, trapping-pattern distribution, and recovery comparison
Figure 3. Trapping incidence across backbones, the distribution of Static, Cycle, and Drift events, and final task performance with generic versus factorized recovery.

GROWING CONTEXT

A current belief, even as history grows.

Figure 4: LOCA-Bench task success from 8K to 256K context across three backbones
Figure 4. At 256K context, PoS exceeds the strongest tested baseline by 10.67-16.00 percentage points across the three backbones. The experiment varies LOCA-Bench environment-description length.

THE COMPUTE TRADEOFF

Better task performance, with additional inference cost.

Belief construction and validation add work even when the task agent spends fewer tokens. Fewer unproductive actions should not be read as lower total compute.

Total tokens
5.06×
1,800.13K vs. 355.70K per episode
Task Agent tokens
−20.9%
281.20K vs. 355.70K per episode

Table 2. RCA-100, Qwen3.7-Plus, mean input + output tokens across 103 cases; compared with Raw Trajectory.

08 GO DEEPER

Video & Resources.

PoS video overview: coming soon. Related work: Graph of States

BibTeX

@misc{luo2026memoryharnessinglonghorizonagents,
  title={Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States},
  author={Yu Luo and Jiamin Jiang and Yimin Zuo and Xidao Wen and Rongchen Gao and Yongqian Sun and Shenglin Zhang and Guiyang Liu and Cheng Zhang and Fang Situ and Qi Zhou and Dan Pei},
  year={2026},
  eprint={2610.01415},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2610.01415},
}