What a larger local model closes, and what card would hold it
Benchmark Report No. 03 / September 2026. Four open-weight models too large for a 12 GB graphics card, on No. 02's seven advisory document tasks against Claude Opus 5. Three of No. 02's four gaps closed to usable with review and parity went from one result to four; filing brief stayed a gap. The best of them, Gemma 4 31B, needs a 32 GB card to hold it whole, by estimate.
Download Benchmark Report No. 03 (PDF, 5 pages, 854 KB) · Builds on Benchmark Report No. 02
- 6 of 7 tasks at parity or usable with review for Gemma 4 31B, the best larger model, against 3 of 7 for the best of No. 02's models
- 1.1 points at most, the change in any score when 50% of Gemma 4 12B was moved into system memory; the wording of 38 of 64 answers changed
The test machine (RTX 5070 12 GB, Ryzen 7 7700X, 63 GB RAM, Windows 11) is a deliberately modest floor that held each larger model partly in system memory; a client deployment sizes the graphics card to hold the whole model, so run times measured here are not reported.
Every model on every task
- Fund term extraction (Field accuracy): Gemma 4 31B 98.6% usable with review; Qwen3.8 27B 99.4% parity; Gemma 4 26B-A4B 99.2% usable with review; Qwen3.6 35B-A3B 98.3% usable with review (point estimate); Best that fits in 12 GB (Gemma 4 12B) 97.8% usable with review (point estimate); Claude Opus 5 100.0%
- Change detection (Material-change recall): Gemma 4 31B 87.0% usable with review (point estimate); Qwen3.8 27B 94.3% usable with review (point estimate); Gemma 4 26B-A4B 89.4% usable with review (point estimate); Qwen3.6 35B-A3B 77.2% gap (point estimate); Best that fits in 12 GB (Ministral 3 14B) 82.1% gap (point estimate); Claude Opus 5 93.5%
- Grounded Q&A with abstention (Answer accuracy): Gemma 4 31B 97.5% usable with review (point estimate); Qwen3.8 27B 97.5% usable with review (point estimate); Gemma 4 26B-A4B 95.0% usable with review (point estimate); Qwen3.6 35B-A3B 93.8% usable with review (point estimate); Best that fits in 12 GB (Qwen3.5 9B) 97.5% usable with review (point estimate); Claude Opus 5 100.0%
- Filing brief (Key-fact coverage): Gemma 4 31B 84.2% gap (point estimate); Qwen3.8 27B 85.0% gap (point estimate); Gemma 4 26B-A4B 75.0% gap; Qwen3.6 35B-A3B 82.5% gap (point estimate); Best that fits in 12 GB (Gemma 4 12B) 76.7% gap; Claude Opus 5 98.3%
- Meeting notes to CRM (CRM accuracy): Gemma 4 31B 91.6% usable with review (point estimate); Qwen3.8 27B 95.5% usable with review (point estimate); Gemma 4 26B-A4B 93.3% usable with review; Qwen3.6 35B-A3B 89.3% gap (point estimate); Best that fits in 12 GB (Ministral 3 14B) 87.6% gap (point estimate); Claude Opus 5 100.0%
- Client drafting (Required-content rate): Gemma 4 31B 99.3% parity; Qwen3.8 27B 99.3% parity; Gemma 4 26B-A4B 98.5% usable with review (point estimate); Qwen3.6 35B-A3B 99.3% parity; Best that fits in 12 GB (Qwen3.5 9B) 98.5% parity (point estimate); Claude Opus 5 100.0%
- Marketing review flagging (Planted-issue recall): Gemma 4 31B 88.9% usable with review (point estimate); Qwen3.8 27B 73.3% gap (point estimate); Gemma 4 26B-A4B 75.6% gap (point estimate); Qwen3.6 35B-A3B 64.4% gap (point estimate); Best that fits in 12 GB (Gemma 4 12B) 66.7% gap (point estimate); Claude Opus 5 86.7%
Each task's primary metric, with the tier under each score. The tiers are No. 02's, fixed before any run: parity is within three points of the frontier with no more unsupported content and no planted trap failed more often; usable with review is within ten points with unsupported content at or under 3%; anything else is a gap. An asterisk marks a tier that rests on its point estimate, because the 95% interval of the gap crosses a tier boundary. Grounded Q&A with abstention is scored on one firm pack of 40 questions, chosen before any larger model ran, for every model including Claude Opus 5, so its figures can differ from No. 02's.
The card under each model is the smallest single graphics card that would hold it whole while it generates: arithmetic from its measured footprint, not tested on that card. gpt-oss:120b, the largest model the selection rule picked, did not load on the test machine (its 56.0 GB expert buffer was more than the 33.9 GB of system memory free) and is not scored.
What it means for a firm
To close No. 02's gaps on a firm's own hardware, start with Gemma 4 31B: parity or usable with review on 6 of 7 tasks, and by estimate a 32 GB card holds it whole (a 24 GB card falls 0.07 GB short once the memory it used while generating is counted). Qwen3.8 27B reached 5 of 7 with more parity results (2 to 1), and a 24 GB card holds it whole. Neither closes the filing brief gap and all three closures rest on point estimates, so briefs still need a full human read and the closed tasks still need a reviewer. Black Lily installs these systems, which is why the study was pre-registered and reports every model on every task; the test that settles it is the same measurement on your own files.
How a deployment works end to end: private on-premise AI.
More from Black Lily Research
All research
Black Lily is an AI implementation consultancy. We install private on-premise AI on hardware you own, for hedge funds, registered investment advisers, and family offices whose best documents are the ones they cannot send to an outside service, and we turn paper, handwriting, and inconsistent spreadsheets into structured data your systems can import. Founded by William Dorman, Founder & CEO.
Phone: (432) 234-3779 · Book a free 30-minute scoping call: cal.com/black-lily/30min
Black Lily home · Hedge funds · RIAs · Family offices · Local AI · Data digitization · Research · Blog · Privacy Policy · Terms of Service