Bart Labs · Field Papers · Research & Benchmarks

We publish what we measure.

Evidence-aware software should be held to evidence. The lab runs studies and benchmarks on its own products — where AI genuinely helps, where it doesn't, and how we know the difference — and writes down the method, not just the conclusion.

Read the paper Get the next one

Can AI tell a good photograph?

Field paper · LrForge

Where AI helps a photographer — and where the human must stay in charge

Abstract. Over more than 3,500 requests against local image-analysis models, we measured where an evidence-aware assistant adds real signal to a photographer's culling and editing decisions — focus, exposure, composition, expression — and where it does not. The finding that shaped LrForge: models are reliable at grounded, checkable observations and unreliable as an arbiter of taste. The product is built on that line. The assistant surfaces evidence; aesthetic authority stays with the person who has to answer for the photograph.

Every claim in the paper traces back to a request in the study set. No cherry-picked demos, no vendor-supplied numbers — the method is in the paper so you can weigh it yourself.

LrForge · 3,500+ requests · ~12 min read

Read the full paper

Headline numbers, with the method attached.

Each benchmark is a short, reproducible claim. Video walkthroughs are being cut for the lab's channel — the numbers below stand on their own until then.

Grounded observation vs. taste arbitration

3,500+ requests measured

Method. Local image-analysis models run across culling and edit decisions on real shoots; each response scored against the frame it described. Split by task type to separate checkable observations from aesthetic judgment.

Evidence capture at bench scale

2.6× faster policy scan, cache on

Method. Internal benchmark on the sensitive-data scan path against a bench-scale corpus, scan cache enabled versus disabled. A measured throughput result on one path, not a whole-system speed claim. The write-up with the harness and corpus is not yet published; ask and we will walk you through it.

Tamper-evidence, independently verifiable

1.3M records sealed into the chain

Method. Every captured interaction is Merkle-linked, and the database rejects edits and deletes on the audit table outright. Verification recomputes the chain and flags any altered record, without needing access to our runtime. The property is structural — a modified record fails verification by construction, not by policy. Boundary: the chain is unkeyed and its summaries live in the database they protect, so it detects accidental corruption and outside tampering, not an insider who already holds database access.

Benchmarks describe measured behavior of Bart Labs products. ForgeShield assists regulated teams in demonstrating and investigating AI use; it does not by itself make an organization compliant.

Want the method behind a claim?

Every number here has a write-up. Ask for the one you care about, or book a walkthrough and we'll run it on your own stack.