Life sciences · Preprint
arXiv · August 17, 2026
Posted before peer review. The findings may change or fail to hold.
This preprint reports a black-box reinforcement learning framework for optimizing agents through complex harnesses, demonstrating empirical improvements on benchmark tasks. The work is a technical contribution describing an infrastructure and algorithmic approach, not a clinical or comparative effectiveness study, and has not undergone peer review.
Preprint. Qwen3-30A3B language model evaluated on agent task benchmarks (ClawGym-Bench, JobBench, OfficeQA).. Intervention: Black-box RL framework using sandbox-based execution, serving proxy, prefix tree trajectory organization, adapted PPO and GRPO algorithms, and mix-harness training..
Pass@1 on ClawGym-Bench improved by 9.98 points through OpenClaw harness with Qwen3-30A3B Pass@1 on ClawGym-Bench improved by 14.81 points through Claude Code harness with Qwen3-30A3B Framework remained stable over 200–400 optimization steps
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed technical report on arXiv describing a machine learning framework and empirical results on benchmark tasks, without peer review or publication in a peer-reviewed venue.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks introduces fundamental challenges. In this work, we present a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses. Concretely, we first build a sandbox-based execution infrastructure that isolates task environments and harnesses within temporary sandboxes for large-scale concurrent rollouts. We then decouple policy optimization from opaque harness execution and place a serving proxy at the model boundary to capture model calls. To reconstruct multi-turn trajectories and improve training efficiency, we organize the captured calls into prefix trees and further adapt both critic-based PPO and critic-free GRPO to optimize over the recovered tree structure. Meanwhile, we maintain training-inference consistency throughout the optimization process. Finally, we introduce mix-harness training, allowing a single model to be jointly optimized by heterogeneous harnesses. With Qwen3-30A3B, black-box RL improves Pass@1 on ClawGym-Bench by 9.98 and 14.81 points through OpenClaw and Claude Code, respectively, while remaining stable over 200-400 optimization steps. Moreover, the framework yields consistent gains on more challenging tasks such as JobBench and OfficeQA. Overall, our framework enables effective, stable, and scalable optimization of general agents through black-box harnesses, supporting unified training across heterogeneous execution systems.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.