What a larger local model closes, and what card would hold it

Benchmark Report No. 03 / September 2026. Four open-weight models too large for a 12 GB graphics card, on No. 02's seven advisory document tasks against Claude Opus 5. Three of No. 02's four gaps closed to usable with review and parity went from one result to four; filing brief stayed a gap. The best of them, Gemma 4 31B, needs a 32 GB card to hold it whole, by estimate.

Download Benchmark Report No. 03 (PDF, 5 pages, 854 KB) · Builds on Benchmark Report No. 02

The test machine (RTX 5070 12 GB, Ryzen 7 7700X, 63 GB RAM, Windows 11) is a deliberately modest floor that held each larger model partly in system memory; a client deployment sizes the graphics card to hold the whole model, so run times measured here are not reported.

Every model on every task

Each task's primary metric, with the tier under each score. The tiers are No. 02's, fixed before any run: parity is within three points of the frontier with no more unsupported content and no planted trap failed more often; usable with review is within ten points with unsupported content at or under 3%; anything else is a gap. An asterisk marks a tier that rests on its point estimate, because the 95% interval of the gap crosses a tier boundary. Grounded Q&A with abstention is scored on one firm pack of 40 questions, chosen before any larger model ran, for every model including Claude Opus 5, so its figures can differ from No. 02's.

The card under each model is the smallest single graphics card that would hold it whole while it generates: arithmetic from its measured footprint, not tested on that card. gpt-oss:120b, the largest model the selection rule picked, did not load on the test machine (its 56.0 GB expert buffer was more than the 33.9 GB of system memory free) and is not scored.

What it means for a firm

To close No. 02's gaps on a firm's own hardware, start with Gemma 4 31B: parity or usable with review on 6 of 7 tasks, and by estimate a 32 GB card holds it whole (a 24 GB card falls 0.07 GB short once the memory it used while generating is counted). Qwen3.8 27B reached 5 of 7 with more parity results (2 to 1), and a 24 GB card holds it whole. Neither closes the filing brief gap and all three closures rest on point estimates, so briefs still need a full human read and the closed tasks still need a reviewer. Black Lily installs these systems, which is why the study was pre-registered and reports every model on every task; the test that settles it is the same measurement on your own files.

How a deployment works end to end: private on-premise AI.

More from Black Lily Research

All research

Black Lily is an AI implementation consultancy. We install private on-premise AI on hardware you own, for hedge funds, registered investment advisers, and family offices whose best documents are the ones they cannot send to an outside service, and we turn paper, handwriting, and inconsistent spreadsheets into structured data your systems can import. Founded by William Dorman, Founder & CEO.

Phone: (432) 234-3779 · Book a free 30-minute scoping call: cal.com/black-lily/30min

Black Lily home · Hedge funds · RIAs · Family offices · Local AI · Data digitization · Research · Blog · Privacy Policy · Terms of Service