UK · SWE/BE → Data/AI/ML · software craftsperson · open source

Where benchmarks and builds live

Lead AI engineer rooted in software and backend systems.

Why these exist

One view of each project. The benchmarks, explorers, and repos stay one click further down.

Charity document extraction

Why
A playgroup needed a fair comparison of models on the same real charity PDFs, with the same fields and the same score.
Problem
UK charity annual reports are PDFs. Eight financial fields have to come out correctly, and models differ in accuracy, failed runs, time, and cost.
Built
A scored extractor across OpenRouter, Doubleword, and V7 Go, plus a public playground of every run.
Result
122 scored runs. The best models reach about 0.95 F1. Provider averages and failures are on the leaderboard, so a choice is a comparison.
Read the project

Python bug benchmark

Why
The Poolside Laguna hackathon asked how laguna-xs.2 behaves on Python bugs next to a field of other models.
Problem
One overall score hides the shape of a model. Easy tasks and hard code-fixes are different jobs.
Built
A 29-model sweep on py-bug-trace, with an explorer, a short write-up, and per-level scorecards.
Result
12th of 29 overall (79%), and 2nd on the hardest level (86.7%): weaker on easy tasks, stronger when the job is fixing code.
Read the project

AgentVetter

Why
Coding agents install skills and tools, then treat that install as trusted code.
Problem
The install is unreviewed code sitting inside a trusted session.
Built
A scan that runs before the agent executes the artifact, and blocks what should not run. The dashboard is the public view of that gate.
Result
A scan-and-block path for Claude Code, built at the Frontline hackathon in London. Open the dashboard to see the gate.
Open the dashboard
Site news · build highlight Live

Model extraction playground — 122 scored runs, rankings & heatmaps.

—
SWE / Backend Data Engineering ML / AI Java Champion Certified AI Engineer 4× Kaggle Expert RAG / LLMs Production Systems Speaker Mentor

Latest snapshot plus historic versions

122

Scored extract runs

12/29

Laguna-XS.2 rank

1.7k★

awesome-ai-ml-dl

Wins & in public

Hackathon outcomes, published benchmarks, and partner shoutouts — each with a LinkedIn post.

Browse builds & pages

Pick a card below — interactive builds open in the browser; static pages are project summaries and guides. Cards are grouped by project family (tag on the left).