ColloidPartner with us

Verified software optimisation · October 2026

Software that proves its own improvements.

From bare metal to a proprietary, machine-optimised stack. Every gain verified before it counts.

Cost per request
−30.9%
Real bugs fixed, same model
+75%
Planted cheats caught
31/31

Why now

Compute is one of the world's largest bills, and much of it is waste.

$723B[1]

Worldwide public cloud spending, forecast for 2025.

29%[2]

Share of IaaS and PaaS spend that organisations estimate is wasted.

~$725B[3]

Capital spending planned for 2026 by Alphabet, Amazon, Meta and Microsoft combined.

How the giants do it

The largest companies already prove that automated optimisation pays.

Google DeepMind · 2025

0.7%[4]

of Google's worldwide compute recovered, on average, by an AI-discovered scheduling heuristic.

Meta · CGO 2019

Up to 8%[5]

faster data-centre applications from a binary optimiser, on top of existing compiler optimisations.

Each win is one technique on one layer, built in-house by a company that can afford a research team. Colloid works across every layer, for everyone.

What Colloid does

It improves real software, and keeps a change only when it is proven.

01

Improve

Colloid finds changes across every layer of a running system: code, data access, runtimes, compilers and operating-system settings.

02

Prove

An independent verifier checks every change and measures its gain with statistical confidence. Where it is decidable, the change is proved formally.

03

Keep

Every verified gain is kept with its evidence: the proprietary raw material for rules, and for the stack we are building.

The vision

From bare metal to our own stack.

Every layer of a computer system is a place to find verified gains. Today the engine works from the operating system up to application code. The destination is a stack of our own, machine-optimised and proven at every layer.

Verified results

Services & applications

Service code, and fixes to real open-source issues.

  • 75% more real GitHub bugs fixed than our first engine generation, with the same AI model.
  • The first fixes on issues created after the model was trained, so no answer could be memorised.
  • Graded by each benchmark's official harness, not by us.

Select a layer to see where it stands. Status as of October 2026.

Verified results

Every number is pre-registered, measured and tested.

−30.9%

cost per request on a production-style web service

95% confidence interval: 27.3% to 34.5% lower. On a hidden holdout workload the search never saw, the gain was 26.6%, and median latency fell 46.9%. Every answer stayed identical, checked independently.

Indexed to the baseline (= 100); lower is better. Cost per request 95% confidence interval: 27.3% to 34.5% lower.

Real-world bugs

Same AI model. 75% more real bugs fixed.

On 30 issues from SWE-bench Verified[6], the industry's standard test of fixing real GitHub bugs, our fourth engine generation resolved 21 against 12. The model, the verifier and the time budget were held fixed.

Exact McNemar p = 0.012, pre-registered before the run.

Issues resolved out of 30 on SWE-bench Verified, one AI model throughout. Gen 4 highlighted.

No memorisation possible

On bugs fixed after the model was trained: from 0 to 7.

Old benchmark issues can leak into a model's training data. So we tested on issues created after the model's release. One change to how the engine searches took it from 0 to 7 of 40. Exact p = 0.016, none lost.

Reported in full: On older issues the same change did not move the score (20 against 21 of 30), within the run-to-run noise we measure: two identical runs differ on 2 of 30. We publish null results too.

Issues resolved out of 40, all created after the AI model was released.

Rigour is the product

Pre-registered, tested and published, failures included.

10

experiments pre-registered before they ran

31/31

planted cheats caught by the verifier

7/7

database optimisations formally proved before running

2/30

our own measured run-to-run noise

Not yet proven, and we say so: That knowledge from one system speeds up the next. Its first test did not pass, and a stronger re-test is designed.

The moat

A growing record of what provably makes software better.

Rarer than code

Each record holds a change, where it applies, and the evidence that it worked, or that it did not.

Private by design

Customer code stays under the customer's control. The evidence and the rules distilled from it stay ours.

Built to compound

Records seed future searches, become rules, and in the end become the stack itself.

Competitors can copy a technique. They cannot copy years of verified evidence.

Who it helps

Anyone whose software runs at scale.

Hyperscalers & big tech

Fleet-wide efficiency at Google scale, where a fraction of a percent moves budgets, and every change ships with proof that it is safe.

Cloud-native companies

Attack the 29% of cloud spend that surveys call waste, without risky rewrites.

AI infrastructure

Verified efficiency for the services around AI workloads. Accelerators are on the roadmap.

Software vendors

Fixes to real bugs, graded by the project's own test suites before anyone reviews them.

Regulated industries

Banking, health and government get an auditable record of correctness and gain for every change.

Platform & SRE teams

Continuous tuning of databases, runtimes and operating-system settings, with the evidence attached.

Milestones

Each step is gated by evidence, not by a calendar.

  1. Achieved · 2026

    • −30.9% cost per request, verified
    • One engine for Python, C, Go and TypeScript systems
    • 75% more real bugs fixed with the same AI model
    • First fixes with no memorisation possible
    • Formal proofs for database optimisations
  2. Next · Months

    • Self-correcting search on long tasks (in progress)
    • Optimising real third-party repositories
    • Rules that work on systems they were not learned on
    • Formal proofs beyond the database
  3. Destination · Year+

    • Our own language and compiler
    • Verification across x86, ARM and accelerators
    • A complete proprietary stack, every component traced to evidence

Colloid

Proof, not promises.

We are looking for investors and design partners that run software at scale and want every efficiency gain delivered with proof.

Get in touch