ALFWorld
88.81%
vs. 72.39%
LongHorizon-HarnessBEYOND MEMORY / LONG-HORIZON AGENTS
Beyond memory. Toward progress.
Remembering what happened is not enough. An agent must understand what has changed, recognize when progress stalls, and recover.
ARXIV PREPRINT / OCTOBER 2026
1Nankai University 2Alibaba Group 3Tsinghua University
* Corresponding author: Yongqian Sun

01 BEYOND RETAINING HISTORY
Large language model (LLM) agents can now undertake increasingly complex tasks, but the way they organize interaction history into memory does not ensure a coherent understanding of the current world. We introduce PoS, an inference-time framework that constructs and continually maintains explicit belief states as the agent's decision context. Each belief combines an estimate of the current world state with unresolved task requirements, making explicit what the agent still needs to learn and accomplish. To keep this belief reliable and actionable, PoS validates its consistency and monitors task progress to detect Belief Trapping, where the agent continues to act without making meaningful progress toward the goal. Recovery is then tailored to both the trapping pattern and the type of unresolved task requirement. Experiments on four benchmarks spanning execution and diagnosis show that PoS achieves the highest overall performance on every benchmark with all three LLM backbones. Ablations demonstrate the importance of consistency validation and recovery, while context-scaling experiments show resilience to context growth. Together, these results support belief construction and continual maintenance as a foundation for long-horizon context management beyond history retention and compression.
FROM STRUCTURED SEARCH TO MEASURED PROGRESSIONGraph of States grounds abductive search in a causal graph and a state machine. PoS maintains a task-conditioned belief for both diagnosis and execution, assesses how that belief evolves, and intervenes when progress stalls.
02 STATE PERCEPTION
Memory retains past interactions. A validated belief describes the current world and what remains unresolved. Progression compares these states over time, so repeated activity can be recognized and redirected.
PoS connects the current belief to the next decision through a repeated update, check, and recovery loop. See the complete method
03 THE COMPLETE LOOP
PoS maintains a current belief, checks its consistency, and measures its evolution across time. The Task Agent selects actions using that belief, an active gap, and recovery constraints when they are needed.

Entities, their states, and their relations, with provenance, confidence, and evidence when available.
A question that remains unresolved, such as distinguishing database waiting from JVM-side processing.
A difference between the current world and the goal, such as making a mug hot before putting it away.
CONSISTENCY COMES FIRST
A candidate update can contain conflicting states or claims unsupported by interaction evidence. The Belief Sentinel returns issues for revision; only the validated belief becomes the next decision context.
The candidate records the microwave as both open and closed. It is not ready to guide the next action.
Schematic update illustrating Section 3.2, not a recorded case step.
04 PROGRESSION ACROSS TIME
PoS checks state updates against evidence before measuring them. It examines what remains unresolved, whether useful changes occur, and whether earlier states recur. Action counts alone cannot tell us when the task needs a different direction.
PROGRESS MONITOR
1 − 0.50 × max(0.50, 0.57) = 0.71
Gap persistence alone does not establish trapping. PoS combines it with stagnation or recurrence.
The mug is cold and inside the microwave.
The mug remains cold.
Diagnosis records progress when this change exceeds 0.30. Confidence values are maintained independently; decreases can contribute too. In execution, the Sentinel assesses whether a transition reduces the gap, acquires needed information, or advances a plausible path.
Compute persistence separately for epistemic and achievement gaps; an empty initial set gives zero. A gap that disappears and later returns is not persistent throughout the window.
Project historical world states onto the same current active gap. Compare aligned Entity, State, and Relation sets over lags 1–4; R is the highest repeat rate. The distance threshold is 0.15. Real belief comparison includes confidence changes.
Section 3.3, Appendix B.2, and Table 8. Detection starts after K = 8 validated transitions and triggers at H ≤ 0.25. The worked examples use aligned symbolic records and the paper's thresholds.
Persistence must coincide with stagnation or recurrence. A maximum over all three signals can mistake unfinished work for trapping; a geometric mean can miss trapping when one signal is absent.
| Situation | (p, S, R) | Geometric mean | Maximum | PoS |
|---|---|---|---|---|
| Stagnation without recurrence | (1, 1, 0) | 1 | 0 | 0 |
| Recurrence despite positive step labels | (1, 0, 1) | 1 | 0 | 0 |
| Progress before gap closure | (1, 0, 0) | 1 | 0 | 1 |
These are illustrative boundary cases, not episode measurements. Lower H indicates stronger trapping signals.
05 TRAPPING-AWARE RECOVERY
Recovery combines two questions: how is the agent stuck, and what remains unresolved? The resulting constraints redirect action selection without supplying the answer or prescribing a tool call.
FACTORIZE. COMPOSE. REASSESS.
Pattern × unresolved requirementThe same task and validated window from Progression.
The relevant world stays unchanged while actions continue.
Avoid repeating the ineffective action under the unchanged belief.
Acquire evidence that distinguishes the remaining explanations.
Once a full window is available, H at or below 0.25 triggers trapping diagnosis.
Recompute health after each validated transition. Keep or update the constraints while trapping persists.
Mechanism illustration of Section 3.3. Static, Cycle, and Drift are tested in order; unmatched cases use generic recovery. Constraints guide action selection while the active gap remains fixed. The actual diagnostic recovery unfolds in the case below.
06 SEE THE DIFFERENCE UNFOLD
Follow the same task through two decision contexts. In PoS, inspect the current requirement, the evidence that changes the belief, and the recovery that restores progress. Raw Trajectory retains the interaction record.
ALFWorld #0096
Reported outcome: 17 actions · task completed
A goal with two requirements.
Put a hot mug in the cabinet.
Temperature and location are separate requirements; completing one does not complete the task.
Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.
Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.
Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.
The task asks for a hot mug in the cabinet.
Both temperature and location matter, even though the instruction is a single sentence.
Action: Receive task
Observation: Put a hot mug in the cabinet.
The baseline puts the mug into the microwave.
Placing the mug in an appliance does not itself heat it.
Action: Put mug in microwave
Observation: The mug is in the microwave.
The baseline moves on without completing the heating requirement.
The action history contains placement, but no completed heating action.
Action: Continue without heating
Observation: The heating requirement has not been fulfilled.
The unheated mug is placed in the target cabinet.
The location now looks right, but the task is still incomplete.
Action: Put mug in cabinet
Observation: The mug reaches the cabinet without being heated.
The baseline examines the cabinet again.
Another look cannot cause the missing temperature change.
Action: Examine cabinet
Observation: The mug remains in the cabinet.
The same cabinet-examination pattern continues.
The context grows while the task-relevant state stays unchanged.
Action: Examine cabinet again
Observation: No heating action is added to the trajectory.
The paper summarizes 36 cabinet examinations in this run.
These repeated observations are activity, not progress toward a hot mug.
Action: Repeat cabinet examination
Observation: 36 cabinet checks in the reported trajectory.
The baseline fails after 50 steps.
The mug reached the right place, but never reached the required state.
Action: End of run
Observation: Failed after 50 steps.
Figures 1–2. Expanded explanation of ALFWorld #0096. Intermediate panels unpack the mechanism, not verbatim action logs. Reported outcomes: Raw Trajectory fails at 50 actions; PoS succeeds in 17.
Put a hot mug in the cabinet.
Temperature and location are separate requirements; completing one does not complete the task.
Current understanding: The task specifies temperature and location requirements.
Unresolved: Make Mug 3 hot, then put it in Cabinet 1.
Guidance: Act on the current state and the remaining requirements.
Mug 3, Microwave 1, and Cabinet 1 become task-relevant entities.
An entity–state–relation view makes the objects and their current relationships explicit.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Make Mug 3 hot, then put it in Cabinet 1.
Guidance: Act on the current state and the remaining requirements.
The belief records what is currently known about the mug and the appliances.
Past actions are evidence for a current state, not a substitute for that state.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Make Mug 3 hot, then put it in Cabinet 1.
Guidance: Act on the current state and the remaining requirements.
The mug needs to become hot and end up in the cabinet.
PoS carries unresolved requirements alongside the world model.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: 1. Make Mug 3 hot. 2. Place the hot mug in Cabinet 1.
Guidance: Act on the current state and the remaining requirements.
The temperature requirement becomes the active gap.
Location alone would not establish the full goal.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Active: make Mug 3 hot. Pending: place it in Cabinet 1.
Guidance: Act on the current state and the remaining requirements.
Further activity leaves the heating requirement unfulfilled.
The display expands the reported stagnant state; it does not invent an exact early tool sequence.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Active: make Mug 3 hot.
Guidance: Act on the current state and the remaining requirements.
The Sentinel reports no consistency issue in the illustrated state.
Validation checks reliability. Progress monitoring asks a different question: is the task moving forward?
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Active: make Mug 3 hot.
Guidance: Act on the current state and the remaining requirements.
The active requirement persists across low-progress interactions.
PoS looks at gap persistence, progress stagnation, and recurrence over a recent window.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Active: make Mug 3 hot.
Guidance: Act on the current state and the remaining requirements.
The reported belief health is 0.1; the missing temperature change remains.
The system diagnoses both how the task is trapped and what kind of requirement remains.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Active: make Mug 3 hot.
Guidance: Act on the current state and the remaining requirements.
Break the ineffective transition and cause a goal-relevant state change.
The constraints guide the Task Agent. They do not replace its choice of action.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Active: make Mug 3 hot.
Guidance: Break the unchanged-state pattern + cause the missing heating transition.
The recovery route uses the microwave to heat the mug.
The active gap stays fixed until supporting evidence establishes that it has been resolved.
Current understanding: Mug 3 is not yet hot and at the counter.
Unresolved: Active: make Mug 3 hot.
Guidance: Address the heating requirement before final placement.
Mug 3 enters Microwave 1.
The world model updates its location; the heating gap is not yet resolved.
Current understanding: Mug 3 is not yet hot and in Microwave 1.
Unresolved: Active: make Mug 3 hot.
Guidance: A location change alone is not evidence of heating.
The selected action makes the mug hot.
A goal-relevant world-state change now supplies a candidate belief update.
Current understanding: Mug 3 is being heated and in Microwave 1.
Unresolved: Active: establish that Mug 3 is hot.
Guidance: Act on the current state and the remaining requirements.
The current belief records the hot mug.
Only after the heating requirement is resolved does placement become the next active requirement.
Current understanding: Mug 3 is hot and in Microwave 1.
Unresolved: Active: put the hot mug in Cabinet 1.
Guidance: Act on the current state and the remaining requirements.
The heated mug is placed in Cabinet 1.
The relation changes while the completed temperature requirement is retained.
Current understanding: Mug 3 is hot and in Cabinet 1.
Unresolved: Check the final hot-mug-in-cabinet state.
Guidance: Act on the current state and the remaining requirements.
PoS completes the reported task in 17 actions.
The final belief establishes temperature and location together, rather than mistaking repeated activity for completion.
Current understanding: Mug 3 is hot and in Cabinet 1.
Unresolved: Resolved: Mug 3 is hot and in Cabinet 1.
Guidance: Act on the current state and the remaining requirements.
Illustrative Raw Trajectory contrast based on Appendix F's diagnostic question. The paper provides no matched baseline trace for t024. These are explanatory interaction summaries, not tool logs or scored results.
This contrast uses the same inventory-service diagnostic question.
The following history illustrates Raw Trajectory; it is not a matched baseline run reported in the paper.
Action: Receive diagnostic request
Observation: Investigate the inventory-service delay.
Slow requests and elevated GC activity enter the interaction history.
These clues alone do not settle which explanation is responsible.
Action: Inspect incident symptoms
Observation: Slow inventory requests coincide with elevated GC activity.
The agent may consider database waiting and JVM-side processing in its reasoning.
Reasoning in the transcript is not an explicitly maintained belief-state object.
Action: Consider explanations
Observation: Database waiting or processing inside the inventory JVM?
A log search adds another interaction to the context.
A longer transcript is useful only when the new evidence resolves something.
Action: Search logs
Observation: The new material does not distinguish the two explanations.
The illustration adds a further search with little new diagnostic value.
Earlier clues remain available, but no structured gap is maintained beside the history.
Action: Continue log investigation
Observation: The diagnostic distinction remains open.
Additional interaction can repeat facts that are already in the context.
The history records the activity without making it a new piece of discriminating evidence.
Action: Inspect further log material
Observation: No additional distinction is established in this schematic.
The same GC clue can be read again without resolving the cause.
Retaining a clue is different from obtaining evidence that separates explanations.
Action: Review earlier observations
Observation: GC remains a clue; the distinction is still open.
A further search adds another interaction.
This contrast illustrates context growth, not a measured baseline outcome.
Action: Search for further log information
Observation: No discriminating evidence is established in this illustration.
Both explanations are still plausible in this illustrative sequence.
Raw history retention alone does not provide PoS's explicit progress checks and factorized recovery.
Action: Review accumulated interactions
Observation: The database-versus-JVM question is still unanswered.
The schematic ends with the diagnostic question unresolved.
No baseline failure rate, final diagnosis, or action count is claimed without its actual trace.
Action: Illustration ends
Observation: No matched baseline outcome is reported here.
Appendix F, p. 26: RCA-100 t024 / Kimi-K3. One recovery episode within a 28-action run. Action indices are zero-based; display frames are not individual tool calls.
Inventory-service requests are slow.
The replay focuses on a recovery episode, not all 28 actions of the reported run.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Earlier observations include elevated garbage-collection activity.
The current belief retains useful evidence without replaying the entire interaction history.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Database waiting and JVM-side processing remain competing explanations.
GC is a clue, not sufficient evidence to decide between every possible cause.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
The next investigation needs evidence that separates the alternatives.
This epistemic gap gives the investigation a specific target.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Further log searches do not settle the distinction.
More text in the context is not necessarily more diagnostic progress.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Progress monitoring uses validated belief transitions.
This panel explains the mechanism; no rejected update is claimed for this case.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Relevant understanding changes little while interactions continue.
PoS can distinguish an unresolved question from merely a long conversation.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
PoS identifies Static Stagnation + Epistemic Gap.
The diagnosis combines an unproductive transition pattern with the kind of missing requirement.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Avoid continuing the same ineffective investigation.
The pattern constraint alone does not specify which evidence would resolve the question.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Avoid repeating the ineffective investigation under the unchanged belief.
Seek evidence that discriminates database waiting from JVM-side processing.
Together, the constraints redirect the agent without handing it a diagnosis or a tool call.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Break ineffective repetition + acquire evidence that separates the alternatives.
The agent selects CPU investigation.
This is an agent-selected response to the constraints, not a tool prescribed by the recovery module.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Recovery initially yields another low-progress interaction.
A useful constraint need not cause an immediate state change.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
A second subsequent action leaves the targeted distinction open.
The active gap is preserved rather than replaced just to make the reasoning appear productive.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
A third low-progress action follows the recovery trigger.
The replay makes the reported delay visible; it does not invent the individual tool outputs.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Subsequent CPU queries show increased activity.
New evidence needs to be interpreted together with the earlier GC observations.
Current understanding: Inventory requests are slow; earlier observations show elevated GC activity.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
Together, CPU and GC evidence strengthen JVM-side processing over database waiting.
CPU activity alone does not uniquely establish memoryPressure; the revision belongs to the accumulated evidence.
Current understanding: CPU + GC observations strengthen the JVM-side explanation.
Unresolved: Distinguish database waiting from JVM-side processing.
Guidance: Seek evidence relevant to the active diagnostic question.
The paper records the epistemic gap as resolved at action 26.
A supported belief change restores progress toward the final report.
Current understanding: CPU and earlier GC evidence support JVM-side processing; the targeted gap is resolved.
Unresolved: Resolved: the targeted database-versus-JVM distinction.
Guidance: Seek evidence relevant to the active diagnostic question.
The agent submits inventory / memoryPressure; both match the benchmark label.
The final result comes from the full run's evidence, not an isolated observation.
Current understanding: Final entity: inventory. Final failure type: memoryPressure.
Unresolved: Resolved: entity and failure type match the benchmark label.
Guidance: Seek evidence relevant to the active diagnostic question.

07 TEST THE IDEA
OVERALL PERFORMANCE
PoS has the highest overall metric among the tested methods in all 12 benchmark-backbone settings. Each comparison below uses the strongest evaluated baseline with the same backbone.
88.81%
vs. 72.39%
LongHorizon-Harness56.38%
vs. 52.57%
LongHorizon-Harness38.83%
vs. 28.16%
PACE / HiAgent45.03%
vs. 41.89%
PACE| Method | ALFWorld | LOCA-Bench | RCA-100 | ClinDiag |
|---|---|---|---|---|
| Raw Trajectory | 62.69 | 43.62 | 24.27 | 38.91 |
| ACON | 66.42 | 49.9 | 27.18 | 39.74 |
| PACE | 67.91 | 17.33 | 28.16 | 41.89 |
| HiAgent | 66.42 | 24.57 | 28.16 | 40.07 |
| LongHorizon-Harness | 72.39 | 52.57 | 26.21 | 40.56 |
| PoS | 88.81 | 56.38 | 38.83 | 45.03 |
| Without consistency validation | 73.88 | 44.57 | 31.07 | 44.54 |
| Without trapping diagnosis | 73.13 | 48.19 | 33.98 | 42.38 |
Table 1. ALFWorld and LOCA-Bench: overall task success. RCA-100: joint entity-and-type accuracy. ClinDiag: overall diagnosis accuracy. All LLM components use the evaluated backbone; methods share benchmark interfaces and environment-action budgets, not identical trajectories. Baseline adaptations are documented in Appendix D.2.
BEYOND BELIEF CONSTRUCTION
Removing either consistency validation or trapping diagnosis reduces the overall metric in every tested benchmark-backbone setting. The comparison below keeps the backbone fixed.
Backbone: Qwen3.7-Plus
PoS88.81%
Without validation73.88%
Without trap diagnosis73.13%
PoS56.38%
Without validation44.57%
Without trap diagnosis48.19%
PoS38.83%
Without validation31.07%
Without trap diagnosis33.98%
PoS45.03%
Without validation44.54%
Without trap diagnosis42.38%
Table 1. Overall benchmark metrics (%); the full method comparison appears above.
With Qwen3.7-Plus, factorized recovery improves final task performance over a generic recovery prompt on all four benchmarks.
Figure 3 (right). These are final outcomes over all evaluated cases, not trap-recovery success rates.

GROWING CONTEXT

THE COMPUTE TRADEOFF
Belief construction and validation add work even when the task agent spends fewer tokens. Fewer unproductive actions should not be read as lower total compute.
Table 2. RCA-100, Qwen3.7-Plus, mean input + output tokens across 103 cases; compared with Raw Trajectory.
08 GO DEEPER
PoS video overview: coming soon. Related work: Graph of States
@misc{luo2026memoryharnessinglonghorizonagents,
title={Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States},
author={Yu Luo and Jiamin Jiang and Yimin Zuo and Xidao Wen and Rongchen Gao and Yongqian Sun and Shenglin Zhang and Guiyang Liu and Cheng Zhang and Fang Situ and Qi Zhou and Dan Pei},
year={2026},
eprint={2610.01415},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.01415},
}