FalseGreen Job

The agent says the change is done. Is it?

FalseGreen independently checks one consequential coding-agent change against the definition of done agreed before the implementation is judged.

$2,999per Job

One bounded consequential engineering scope. Production Beta.

Independent acceptance

Tests passed
The agent said done
FalseGreen: FAILED

The agent doesn’t get the final say.

When to buy

Use FalseGreen when being wrong costs more than the Job.

Examples of consequential agent-written changes:

Authentication / authorization

Billing / payments

Database migrations

Infrastructure / Terraform

Access control

Data integrity

Production recovery

Release-critical backend behavior

Not worth a Job:

Lint cleanup

Cosmetic refactors

Routine changes adequately covered by ordinary engineering controls

“We just want another AI reviewer”

If this change is commodity work, don’t buy a FalseGreen Job.

The acceptance gap

Green evidence. Wrong result.

Evidence can be green

  • Tests passed
  • CI passed
  • Static analysis passed
  • Agent self-review passed
  • Human review passed

The Job can still be wrong

  • Required behavior missing
  • Important invariant violated
  • Edge case outside the test suite
  • Wrong production state
  • Agent declared completion too early

Tests and review produce evidence. They do not define what “done” means.

How a Job works

Freeze done before judging the work.

  1. 1. Define done

    A human defines the consequential outcome.

  2. 2. Freeze it

    FalseGreen freezes that definition before implementation is judged.

  3. 3. Build

    The coding agent implements and repairs normally.

  4. 4. Verify

    FalseGreen independently evaluates the exact resulting source/state.

  5. 5. Result

    Accepted / Failed / Insufficient Evidence.

If repair is allowed, the agent repairs against the same frozen boundary. The target does not move.

How to define a good Job →

Proof

FalseGreen has actually said FAILED.

Even when ordinary signals looked good.

Case study · FalseGreen on FalseGreen

Tests passed. FalseGreen still refused to call it done.

FalseGreen found a production integrity/authority problem that made the work unacceptable — after all the ordinary signals were green.

Different parts of the system disagreed about what was authorized. FalseGreen said FAILED.

Read the case study →

Case study · Ada / SPARK

The implementation looked correct. The evidence said otherwise.

Cross-unit formal evidence showed the implementation could appear correct while the frozen boundary was not established. The agent repaired against the same frozen boundary until the evidence supported ACCEPT.

2h 44m from frozen target to FalseGreen ACCEPT. Less than 30 minutes of estimated human attention.

Read the case study →

See the evidence behind independent acceptance →

Qualification

Is this a FalseGreen Job?

Tell us the change. We can determine whether it belongs inside one bounded Job.

Probably yes

  • A consequential change is happening now
  • A coding agent is implementing or repairing it
  • Being wrong would materially hurt
  • You can define what must be true
  • You want an independent result bound to the exact work

Probably no

  • Commodity refactor
  • Style / lint cleanup
  • Low-consequence internal change
  • You only want another PR reviewer
  • Being wrong would not justify $2,999

The offer

One FalseGreen Job

$2,999

One bounded consequential engineering scope under one frozen definition of done.

  • Definition of done frozen before judgment
  • Independent verification of the exact resulting work
  • Accepted / Failed / Insufficient Evidence
  • Eligible repair + reverification against the same boundary
  • Durable source-bound report

Support today: Python · Rust / Anchor · Go · Ruby / Rails · Ada / SPARK · Fortran · Scoped Terraform / AWS

Trust

FalseGreen does not promise bug-free software.

  • Result is bounded to the frozen Job
  • Result is bound to the exact work that earned it
  • Insufficient Evidence is allowed
  • Customer tests/CI/review remain useful evidence

Trust & Security →

What consequential change is your agent shipping next?

If being wrong about it would cost substantially more than$2,999, it may be a FalseGreen Job.