Software engineering / EvaluationSEPTEMBER 2026

What we tested in ReproLab /

The tests behind ReproLab: inputs, actions, observed results and the limits of that evidence.

Reproduce a bug in an isolated Node execution environment, review a bounded repair and verify it against the original tests.

Internal local testing, using synthetic inputs and owned assets. The results below establish the observed software behavior in those scenarios. They do not establish customer ROI, uptime or external deployment readiness.

01

An actual failing baseline and passing candidate

The HTTP execution suite passed 14 ReproLab checks. Its synthetic sum implementation returned zero for one plus one; the original regression command exited1. Replacing subtraction with addition produced an actual exit 0 verification against the unchanged test. The package retained both results. These observations establish the fixture's repair, not general defect-detection accuracy.

02

The same sequence through the browser

A separate headed browser pass registered a workspace, imported the folder, created the report, ran the baseline, staged sum.js, reviewed its diff and verified it. The downloaded ZIP contained matching evidence and patch content. Reload retained the verified run. Desktop 1440 px and mobile 390 px / 320 px layouts were visually inspected, including the candidate view.

03

Patch boundaries and an applicable export

The API rejected test-file and package-manifest patches, new-file additions and traversal paths in the bounded acceptance scenarios. A separately exported diff passed actual git apply --check against the original clean fixture. That check did not modify the repository and is narrower than testing every possible multi-file diff or newline combination.

04

Recovery and model evidence stay separate

Four isolated recovery checks recorded one Docker run across failed delivery and restart, with identical recovered output. One consented Gemini extraction also proposed the correct source-only addition patch, while retaining uncertainty about other behavior. That AI result was not automatically executed; the manual execution evidence and the model proposal remain separate claims.

What this evidence does not establish

Supported imports are dependency-free Node 22 UTF-8 projects, bounded to 100 files,30 KB per file and200,000 total characters in the inspected interface.

The manual patch path targets existing non-test source files. It does not install packages, modify test expectations or merge a remote branch.

Passing the supplied regression suite is evidence for those tests, not proof that every behavior is correct.

Model requests require a configured provider and the application's consent gate. Confidential customer code was not used in the recorded model test.

The forced recovery test covered completed evidence awaiting delivery; it did not force every lease-expiry, mid-command crash or production disaster scenario.

APPLY THE THINKING

A similar problem in your business?

Bring a real workflow, representative inputs and the result that needs to be reliable.

Shape a project brief