R&D
R&D

Hand it a bug, get back a mergeable fix.

Back to Solutions
1h / PR5min / PR
·100 PRs in parallel

A customer reports a bug on the ordering page: oat milk is still priced as dairy milk, and several customers have asked about it. This kind of ticket used to sit in an engineer's queue — reproduce, locate, assess risk. Now it goes straight to a R&D avatar, running entirely on the engineer's own Karma Box — reading and writing code never leaves that machine.

Root cause located

10:20 — the owner flags the bug in chat. 10:34 — the avatar replies with the root cause already located: price calculation reads the product's default spec, not the customer's actual selection — cart/price.ts:142. It writes a failing test to reproduce the bug first, tagged "running on the engineer's own Karma Box — reading and writing code needs a local body" — not a line of production code touched yet.

R&D avatar locating the root cause and writing a failing test to reproduce it

Red to green — then it asks before merging

The avatar writes against the current codebase, never from memory of an API it half-remembers. The fix is one line: look up the price by the customer's chosen spec ID, falling back to the default only when none is selected. Before: oat milk should cost ¥31 but charges ¥28 — 1 test failing. After: 3 tests pass, including a negative case that guards specifically against this regression.

Before/after test results and the merge request waiting for approval
3new test cases added, including 1 negative regression case — merge request #142 touches just 1 file

Merging and shipping always wait for your nod — what it hands back is a mergeable fix, not a suggestion.

Real product UI: the R&D Factory

The walkthrough above is a staged scenario; in the real product this capability lives in the "R&D Factory" — a team of engineering avatars working against your actual code repository, able to read and write code, run tests, open merge requests, and show every step's progress and output in one panel.

Real product screenshot of the R&D FactoryR&D Factory task execution and code output view

Why you can't just hand a model your repo and walk away

Paste a bug description into a generic chat window and a model will hand back a patch that looks plausible — but it's written from a half-remembered impression of an API in its training data, not from your codebase as it exists right now. It doesn't know cart/price.ts was touched this week, doesn't know specs went from an array to a Map last month, and won't run your existing tests before it writes a line. This is exactly what the R&D Factory is built to prevent: the avatar must read the repo's current state before writing anything, must write a failing test that reproduces the bug before touching the fix, and must watch that test turn green before calling the job done — never just assert that it probably works. The order is fixed on purpose. Skip any one step and the fix can be wrong while still looking right.

Four steps to put it in your development workflow

1. Connect the real repo, not just "permissions." Point the R&D Factory at the repository you actually work in (not an exported snapshot) so it always reads the latest commit — but in the beginning, grant it merge-request authoring rights only, never direct-to-main merge rights. Merging always stays a human click. 2. Start with one low-risk bug category. Pricing calculation errors, copy typos, stale config values — contained, easy-to-verify tasks that let you build trust in its output before handing it anything more consequential. 3. Require test-first, always. Whether or not you use the Factory's built-in flow, insist it write a failing test that reproduces the bug before it touches the fix — that's the most reliable evidence a fix actually works, not its own description of what it did. 4. Treat every merge request as something to review, not a final answer. Review it like a colleague's PR: is the diff scoped to what was asked, did it introduce new dependencies, does the new test actually cover the original bug.

How to tell whether this is actually saving time

Don't count it as working just because the avatar replied. Track more specific signals: time from bug report to a mergeable fix; the rate at which merge requests get sent back for rework (a high rate means it doesn't yet understand enough of the codebase — scope it to smaller tasks); the share of new tests that are negative/regression cases, not just happy-path checks; and whether your actual review time is shorter than fixing it yourself from scratch. If every merge request needs a line-by-line rewrite, the workflow has just moved the effort around, not reduced it.

Common mistakes

The most common mistake is treating "tests pass" as sufficient proof a fix is correct — the test itself can be wrong, or cover only the one case mentioned in the bug report while missing an adjacent edge case. A second is granting auto-merge too early: once merge request volume climbs, teams relax line-by-line review because things have "looked right" for a while — that's exactly when risk starts accumulating quietly. A third is treating the R&D Factory as a replacement for code review rather than an accelerant for the prep work before it — it can make locating, reproducing, and fixing fast, but whether a change fits where the product is actually headed still needs a person to decide.

Frequently asked questions

How large a change can it handle? Start with small, independently verifiable changes — a single-file bug fix, a local refactor, filling in missing test coverage. Changes that cross multiple modules or involve architectural decisions should be broken into smaller tasks by a person before being assigned. Does code leave the machine? The R&D Factory can run entirely on your own Karma Box — reading and writing code never leaves that machine. Whether to route non-sensitive steps through a cloud model is your call. If a fix turns out wrong, how do you trace and roll it back? Every merge request carries a full change record and its corresponding test results, so you can pinpoint exactly which merge introduced a problem — it goes through the same version control and rollback process as any human-authored commit, with no special-cased shortcut.

How one team actually adopted it

One 6-person team building a mini-program spent its first week doing exactly one thing: connecting the R&D Factory to its repo with permissions set to "can open merge requests, cannot merge," then assigning it 11 low-priority bugs that had been sitting untouched for a month. Week one: 9 merge requests came back, 7 passed on the first try, 2 got sent back for insufficient test coverage. Week two: the team fed those 2 rework cases back in as reference, and started assigning medium-complexity tasks — the kind touching two files and requiring a bit of business-logic understanding. By week four, the team's own measurement showed these small fixes had gone from "whenever there's time" to "same-day turnaround," and the freed-up time went straight into performance work that had been permanently deprioritized. Not one merge happened automatically anywhere in this process — the tech lead spent a fixed 20 minutes a day reviewing that day's merge requests, and those 20 minutes turned out to add up to far less than handling the same bug volume the old way, spread thin across the whole week.

Before connecting the R&D Factory to any code path touching production, payments, or user data, validate fix quality, merge-request hygiene, and your team's actual review cadence on non-critical modules first, then expand scope gradually.

Start building your avatar — free

Other industry solutions

Legal

Not a suggestion — an edit you can accept.

Customer Service

It has its own address.

Marketing

One piece of content, native to every platform.