My view on Recursive Self-Improvement (RSI) is this: general long-horizon capability is a prerequisite for RSI.
AI modifying its own code is not the hardest part to imagine. What matters is that a round of work gives it better tools, more reliable execution, and more useful experience—and that these improvements participate in the next round of work, helping it improve itself again.
Intelligence is not just being invoked. It begins to compound.
Here I mean an agent system continually improving its own code, tools, and ways of working, not training model weights. This need not begin with a fully autonomous superintelligence. It can begin in a real project: people define direction, provide feedback, and accept results; agents continuously help improve their own operating environment.
LoopX is my practice along that path. It is both the long-horizon agent meta-harness I use to advance projects and the system that LoopX agents continuously develop, validate, and maintain.
Why “general long-horizon” comes first
“Recursive” means that the next improvement builds on the previous result.
If every round requires explaining the goal again, rebuilding the environment, and finding the last conclusion, a system may repeatedly produce changes without accumulating capability. One excellent inference does not necessarily become a system that keeps getting better.
Long-horizon capability carries goals, work state, validation results, and experience across contexts, sessions, executors, and system versions. A different agent can take over the same work. After an environment restart, the improvement can still become the next starting point.
LoopX aims to support many sustained, long-running sessions, giving agents enough time to research, execute, validate, and correct their work. Externalized state supports those sessions and preserves goals, constraints, and results through context compaction, interruptions, or handoffs, so the work already advanced can continue to accumulate.
“General” matters too. Real self-improvement does not stay inside one fixed task type. Today the system might discover a tool defect; tomorrow it might improve a research method; later it might change how agents collaborate, then evaluate what these changes do to overall output. It must connect research, implementation, evaluation, and adoption—not merely repeat a specialized scoring script.
The prerequisite I mean is therefore not “being online long enough.” It is this: the system can keep advancing open-ended goals and carry its improvements into subsequent work.
From using AI on a project to improving the system with an improved system
Models provide intelligence, runtimes provide tools and execution environments, and a long-horizon system connects individual executions into ongoing work.
In this structure, the object of self-improvement is broader than model weights. Tools, context organization, workflows, validation, memory, and human–agent collaboration can all be improved.
Sakana AI’s Darwin Gödel Machine demonstrates a related research direction: agents modify their own code, retain candidates through task evaluation, and continue exploring from existing candidates. It strengthens my conviction that system-level evolution deserves serious attention.
Anthropic’s work on harnesses for long-running agents approaches a complementary problem: with finite context, incremental execution and resumable artifacts carry work across sessions.
My interpretation is that these directions converge: self-improvement produces changes; general long-horizon capability lets them accumulate. One creates the next step. The other lets the system actually climb it.
LoopX treated itself as a long-term task from the beginning
This history should not begin with the last few weeks.
Early: meta-bootstrapping in goal-harness. On May 31, 2026, the first public commit established a small control plane named goal-harness. On June 1, the repository already added a Goal Harness Meta Goal: treating its own development as a goal, checking runnability, continuation, and validation, then advancing the next improvement.
Bootstrapping was not a feature added after the project matured. Very early, it was using itself to manage its own development. As new problems emerged, goal state, successor work, and validation mechanisms grew through that same process.
By June 19, the public repository had a dedicated repository-level self-iteration case. The case’s fixed June 20 snapshot recorded 801 commits and described parallel development across the control plane, planning, validation, interfaces, and multiple workstreams—not one isolated feature. The project had not yet officially become LoopX.
Middle: from a bootstrapping tool to a general long-horizon system. On June 21, goal-harness was renamed LoopX. Work expanded into multi-agent, cross-runtime, and community collaboration. July’s peer runtime and v0.2 control plane developed work ownership, execution boundaries, and handoff into product contracts.
The important change was not adding more models. It was enabling different executors to advance work around the same durable state. Implementation, review, and validation could belong to different roles, while the work still needed to be adopted and accepted.
Later: from an engineering control plane to a personal workspace and community evolution. On September 6, LoopX 1.0 introduced Personal Workspace. Beyond the state kernel, work gained a more complete everyday interface. People could participate in goals, decisions, and results through the workspace and IM; agents continued improving the underlying system, tools, and product surfaces. System bootstrapping and community contributions became intertwined. LoopX was no longer just one person’s development loop.
Together, these stages tell the self-iteration story I want to share: first, make the system itself a long-term goal; then make long-horizon work a general capability; finally, let more people and agents keep evolving the same product.
Four months of public development
I divided the same public Git history at those two milestones and consistently excluded merge commits:
| Stage | Date range | Non-merge commits | Daily average | 00:00–08:00 commits | Overnight share |
|---|---|---|---|---|---|
| Early · Meta bootstrapping | May 31–June 20 | 739 | 35.2 | 298 | 40.3% |
| Middle · General long-horizon system | June 21–September 5 | 3,870 | 50.3 | 1,399 | 36.1% |
| Later · Workspace and community | September 6–October 1 | 2,239 | 86.1 | 556 | 24.8% |
The fixed public-history snapshot contains 8,613 commits, of which 6,848 are non-merge commits. The three stages above sum to those 6,848. All times use Beijing time.
In the latest 14 complete calendar days, September 18–October 1, the daily average reached 96.9 non-merge commits: 1,356 in total, with a daily peak of 144. 384 commits, or 28.3%, fell between midnight and 08:00. All 14 days had overnight commits.
The lesson is not that the overnight share must keep increasing. In fact, it declined across the stages while daily activity rose. More useful questions are whether work continues across day and night, whether improvements enter the next operating environment, and whether more people and agents sustain useful output.
These counts describe development across the whole project, including community contributions. They are not a “fully automated RSI” percentage or a measure of capability improvement. They make the scale and timing of ongoing development auditable. The core self-iteration evidence remains the system being used in real work, receiving feedback, being improved, and then being used again.
Long-horizon systems do more than “run a little longer”
Starting from externalized goals and task state, LoopX has progressively organized continuation, authorization, resource budgets, validation, and accumulated experience into one control plane, with Personal Workspace and IM bringing people into the process. Models and runtimes can differ; the work’s goals, constraints, and acceptance evidence must continue. See the public architecture and product entry points.
Three capabilities matter most to me:
Goal continuity. Completing a local task does not mean the overall goal has been achieved. The system must replan from acceptance results and keep moving toward the actual goal.
Accumulating feedback. User corrections, effective improvements, and failed attempts should not live only in a single conversation. They need to change the environment of subsequent execution.
Carrying improvements forward. Research cannot stop in a document, and code changes cannot stop on a branch. They must be validated, adopted, and tested again in real work.
Connecting these creates a chance to turn more time, compute, and parallel agents into accumulated capability—not just more independent outputs.
I regard LoopX’s human-in-the-loop bootstrapping in real engineering as an early form of system-level RSI. The point is not to rename ordinary development. It is to make the system doing the work an object of continuous improvement, then use that improved system in the next round.
Humans in the loop do not negate recursion
I do not think early RSI must begin by excluding humans entirely.
High-quality human reward helps a system identify what is worth doing and what actually works. People provide direction and key judgments; the system handles sustained execution and continuation. Feedback need not become another explanation from scratch. It can become part of the environment, work state, and validation requirements.
Recursion depends on whether the improved system participates in subsequent improvement, not on whether a person speaks in every round.
At the same time, candidate changes cannot earn “success” by relaxing their own evaluation standards. Real adoption, independent validation, and reversible work stages are foundations for a feedback loop that accumulates value.
Where I want to take this
Models are becoming stronger, tokens cheaper, and the number of agents one person can organize is growing.
The next frontier is not only completing more tasks. It is a digital team gradually developing better methods through long-term work: better tools, more effective collaboration, more rigorous validation, and experience that later work actually uses.
One-shot intelligence solves a problem. Long-horizon intelligence accumulates the ability to solve problems.
That is my ambition for LoopX: one person defines goals and provides key judgments, while a digital team keeps advancing. It not only completes work; it improves the system that completes the work.
Maximize useful agent output. Minimize human attention. General long-horizon capability is the foundation of that path—and, in my view, a prerequisite RSI cannot bypass.
Data and further reading
The statistics are fixed at b9b758c, the last main-line commit before the end of October 1, 2026 in Beijing time. Calendar days use Git committer timestamps in Asia/Shanghai; overnight means 00:00 inclusive to 08:00 exclusive. We count all commits reachable from that snapshot, excluding multiple-parent commits (merges) from the headline measure. The stages follow the first public commit, the LoopX rename, and the 1.0 workspace release. Daily-average denominators are 21, 77, and 26 calendar days; the first day is the partial day on which the public history began. Another 1,765 merge commits are excluded from the 6,848 headline. Commit objects are not independent features or accepted outcomes, and timestamps do not identify whether authors used AI.
For public adoption and subsequent corrections, see the RSI-Harness adoption audit, PR #4984, and follow-up correction #5119. These show how improvements enter actual paths; their implementation details are not the main narrative here.
Project: LoopX.
Download aggregate statistics and method (JSON) · Stages use calendar dates, not the exact hour of each milestone.
Check the fixed history in a full Git clone
git rev-list --count b9b758ca4dfc30d0d610b45398fc31acd62781f6
git rev-list --count --no-merges b9b758ca4dfc30d0d610b45398fc31acd62781f6
TZ=Asia/Shanghai git log --no-merges b9b758ca4dfc30d0d610b45398fc31acd62781f6 \
--format='%cI' # committer timestamps; group by the stated calendar intervals