Papers
arxiv:2610.00948

GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution

Published on Oct 1
· Submitted by
Zhongxiang Dai
on Oct 7
Authors:
,
,
,
,
,
,

Abstract

The executable harness surrounding a GUI model determines how observations are assembled, actions are executed, and verification, recovery, and termination are controlled. Compared with harness optimization for non-GUI agents, automatically optimizing this harness poses three coupled challenges: reconciling model intent with observed visual effects, diagnosing failures under variable execution outcomes, and identifying recurrent failure patterns across tasks and translating them into reusable runtime changes. We introduce GUI-HARVEST, an automatic harness optimizer that enables self-improving GUI agents with frozen backbone models. First, to ground diagnosis in observed action effects, it aligns model outputs and executed actions with before-and-after screenshots, tying findings to specific interface transitions. Second, to account for execution variability, it treats repeated runs of the same task as a joint evidence unit, using within-task comparisons to locate outcome-relevant behavioral differences. Third, it consolidates verified findings across tasks into recurring failure patterns, maps them to bounded source-code edits with predictions recorded before evaluation, and checks the predicted behavioral effects alongside task performance through repeated execution. Experiments on OSWorld-Verified show consistent held-out gains across six general-purpose open, GUI-specialized open, and proprietary backbone models; Qwen3-VL-32B-Instruct gains 12.33 points on the full suite. Frozen-harness transfer improves GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps without further optimization. With the same backbone and initial harness, GUI-HARVEST outperforms Self-Harness and Meta-Harness, suggesting that GUI-specific diagnosis and validation help harness improvements generalize to unseen tasks. The code is available at https://github.com/GaryYang12345/GUI-HARVEST.

Community

Paper submitter

👋 We’re excited to share GUI-HARVEST, which enables GUI agents to improve automatically by evolving their execution harness while keeping model weights frozen.

The idea: learn from what actually happens on screen. GUI-HARVEST compares screenshots and action traces across repeated runs, identifies recurring failures across tasks, and translates them into reusable code changes. Each change must pass checks on both task performance and its predicted behavioral effects.

📊 The figure highlights results on OSWorld-Verified:

  • Six backbones, three step budgets: GUI-HARVEST achieves the highest score in every model–budget comparison shown.
  • Optimize at 15 steps, evaluate at 15/50/100: the optimized harnesses remain frozen across evaluation budgets.
  • Strong gains without weight updates: Qwen3-VL-32B improves by 12.33 percentage points over its initial harness at 15 steps, while Gemini 3.1 Pro reaches 79.14% at 100 steps.

We hope this helps make GUI agents more capable through systematic learning from execution experience. Feedback and discussion are welcome!

💻 Code: https://github.com/GaryYang12345/GUI-HARVEST
📄 Paper: https://arxiv.org/abs/2610.00948

radar

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.00948
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.00948 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.00948 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.00948 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.