Bart Labs · Field Papers · Research & Benchmarks
Evidence-aware software should be held to evidence. The lab runs studies and benchmarks on its own products — where AI genuinely helps, where it doesn't, and how we know the difference — and writes down the method, not just the conclusion.
Published · Photography
Field paper · LrForge
Abstract. Over more than 3,500 requests against local image-analysis models, we measured where an evidence-aware assistant adds real signal to a photographer's culling and editing decisions — focus, exposure, composition, expression — and where it does not. The finding that shaped LrForge: models are reliable at grounded, checkable observations and unreliable as an arbiter of taste. The product is built on that line. The assistant surfaces evidence; aesthetic authority stays with the person who has to answer for the photograph.
Every claim in the paper traces back to a request in the study set. No cherry-picked demos, no vendor-supplied numbers — the method is in the paper so you can weigh it yourself.
One evidence-heavy field paper when we have something worth publishing. No weekly AI roundup.
Benchmarks
Each benchmark is a short, reproducible claim. Video walkthroughs are being cut for the lab's channel — the numbers below stand on their own until then.
Benchmarks describe measured behavior of Bart Labs products. ForgeShield assists regulated teams in demonstrating and investigating AI use; it does not by itself make an organization compliant.
Every number here has a write-up. Ask for the one you care about, or book a walkthrough and we'll run it on your own stack.