← The Business DiagnosticWhere this sits. This is the Gate Zero verdict — the go/no-go stage that follows the Business Diagnostic, and only when the diagnosis concludes something should be built.Illustrative sample. “Arvela Logistics” is a fictional company and every number below is invented. The method, the scoring, and the format are exactly what you receive — built on your numbers instead.
A sample Gate Zero Readout.
Page 1 · Decision memo
Gate Zero Readout — Arvela Logistics · March 2026 · Prepared by Luca-Ștefan Tamaș, LT Strategy Partners
Ship the shipment-document assistant. Consolidate the knowledge base before any customer chatbot. Don't build the pricing engine.
Of the three use cases assessed, one is ready to build now: an assistant that reads booking emails and attached shipment documents and drafts structured entries for the operations team to confirm. The other two fail Gate Zero today — one temporarily, one on design.
- The document assistant clears the gate. Five operations staff spend roughly nine hours a week each re-keying data from PDFs and emails — about €54,000 a year at a conservative base case, using Arvela's own loaded-cost figures. The grounding content exists, is accessible today, and every output is confirmed by a human before it enters the TMS — so the error budget is honest.
- The support chatbot is “not yet”, not “no”. The answers customers need live in seventeen inboxes and a wiki last updated fourteen months ago. Built today, the assistant would confidently repeat stale answers. The unlock is organizational, not technical — and it's named on page 4.
- The pricing engine fails on stakes, not on ambition.A wrong extraction gets caught by the operator who confirms it. A wrong price goes straight to a customer. With spot-quote history living in email threads, the error rate cannot be validated to the tolerance the workflow demands — this is where AI won't pay off for Arvela as scoped.
What leadership already suspected, confirmed with numbers: the re-keying burden your operations lead has flagged for two years is the strongest AI case in the company — stronger than either of the two ideas that came from vendor conversations.
Owner & 90-day stepOperations director. Scope the document assistant against the gate checklist on page 3; in parallel, assign one owner to the knowledge base with a 90-day consolidation target.
Page 2 · The opportunity portfolio
Three use cases, one map.
Value at stake →Production viability →Not yet — close the gap first
Build now
Don't build — and here's why
Quick wins, modest stakes
A · Document assistant AI Act: minimal
B · Support chatbot AI Act: limited
C · Pricing engine AI Act: minimal margin-critical
Dimension scores at walkthrough depth
- G On footing — no blocker found at walkthrough depth
- A Addressable gap — named condition must be met
- R Gating problem — blocks a build verdict as scoped
Map badges show each use case's provisional EU AI Act tier. Additional flags — like margin-critical on C — mark business stakes; they are not AI Act tiers.
- A / Validation (amber): extraction accuracy needs a 200-document ground-truth set before launch — one week of operator time, scoped in the gate checklist.
- A / Exposure (amber): shipping documents carry personal data; processing must stay in the EU, and the extraction step should run self-hosted or in an EU region with no third-country transfer.
- B / Value (amber): the interruption cost is real but unmeasured — the base case needs two weeks of ticket counting before it can be banked.
- B / Data reality (red): no single, current source of the answers exists — they're scattered across seventeen inboxes and a wiki untouched for fourteen months.
- B / Grounding & retrieval (red): fragmented, stale content is precisely what retrieval tuning cannot rescue — the score follows the anchor until the knowledge base is consolidated.
- B / Validation (amber): correctness is checkable by sampling against the ops team — but only once a consolidated source defines what a correct answer is.
- B / Cost & ownership (amber): query volumes and per-answer economics can't be modeled until the assistant has a source to answer from, and no owner is named yet.
- B / Exposure (amber): customer-facing, so EU AI Act transparency obligations apply, and transcripts would carry customer personal data — both addressable at design time.
- C / Value (amber): the margin upside is plausible but can't be validated from unstructured quote history — value that can't be measured can't be banked.
- C / Data reality (red): spot-quote history lives in email threads; no structured record exists to train or validate against.
- C / Grounding & retrieval (red): the pricing context that would ground a quote — lane, season, client history — isn't retrievable from any current system.
- C / Error tolerance (red): a wrong price reaches a customer with no human between the model and the harm — the workflow as designed has no error budget.
- C / Validation (red): with no ground truth, the error rate cannot even be measured to the tolerance the workflow demands.
- C / Cost & ownership (amber): per-quote economics look fine on paper, but they're meaningless while the error rate is unmeasurable.
The binding constraint at Arvela: data consolidation, not ambition, gates everything after use case A.
Page 3 · Top use case — brief & formal gate
A · The shipment-document assistant.
- Business process
- Inbound booking requests: email + attached PDFs (bills of lading, packing lists, customs forms) re-keyed into the TMS by the operations team.
- Impact, conservative base case
- 5 operators × ~9 hrs/week re-keying × Arvela's loaded cost ≈ €54,000/year, before error-correction time. Assumption declared: volumes from the ops lead's estimate, not measured logs.
- Buy / integrate / build
- Integrate: an extraction model behind a review UI, wired to the existing TMS API. No custom model training warranted at this volume.
- Data readiness
- Green at walkthrough depth — documents accessible in the shared mailbox today; ground-truth set needed pre-launch (see gate).
- Cost band & time to production
- Low five figures, 8–12 weeks to a production system with validation, monitoring, and handover. Range, not a quote — a fixed price follows a scoped audit.
- Success metric
- Hours of re-keying per 100 shipments, measured before and ~30 days after launch.
The gate
Build now — subject to the checklist below.
- Capability target: ≥95% field-level extraction accuracy on the 200-document ground-truth set before go-live; every output human-confirmed until then and sampled after.
- Safety target: no auto-commit to the TMS — operator confirmation stays in the loop until three consecutive weeks inside the error budget.
- Cost at scale: unit cost per processed shipment stays under the agreed ceiling at 2× current volume; measured, not assumed.
- Sign-off domains: operations owner · security/data-protection · engineering · the operators who use it.
- Walk-away conditions: if the ground-truth set can't reach the accuracy floor, or EU-only processing can't be met at acceptable cost, stop — the readout's verdict flips to “not yet” and says why.
- What would change this verdict: a TMS migration mid-build, or volumes 3× below the estimate used in the base case.
Confidence: scores marked at walkthrough depth only — verified against described systems, not inspected pipelines. The paid audit tests these against the actual mailbox and TMS before any build.
Page 4 · Where AI won't pay off for Arvela
The honest no's — and what would change them.
Not yet B · Customer support assistant
The value is real — status queries interrupt the same team that re-keys documents. But the grounding content fails Gate Zero: the answers live in seventeen inboxes and a wiki last touched fourteen months ago. An assistant built on that would answer confidently and wrongly, in front of customers, under the EU AI Act's transparency obligations for customer-facing bots.The unlock: one owner, one consolidated knowledge base, a 90-day target — then this use case re-enters the gate with a realistic path to “build”.
Don't build C · Dynamic pricing engine
Not because the model can't produce prices — because the workflow has no error budget and no validation path. A wrong extraction is caught by an operator; a wrong price reaches a customer directly and costs margin or trust. The spot-quote history that would train and validate it lives in unstructured email threads, so the error rate can't even be measured to the tolerance required.The cheaper answer that survives scrutiny: a rules-based margin floor with quarterly review — nearly free, no surprise behavior. What would change the verdict:two years of structured quote/outcome data and a human approval step that doesn't erase the speed advantage.
Page 5 · Gaps, next steps, and the map forward
What happens next — with us or without us.
- Assign the knowledge-base owner (yours to do, this week). No engagement needed — one name, one 90-day consolidation target, and use case B stops being blocked.
- Build the ground-truth set (yours to do, ~1 week of operator time). 200 recent shipments, fields verified by hand. Whoever builds the assistant will need it — us or anyone else.
- Scope the document assistant against the page-3 gate. The open questions — actual extraction accuracy on your documents, EU-only processing cost — are exactly what a fixed-scope review of the candidate against these gate conditions tests, on your real documents.
- Re-gate the chatbot in 90 days. If the knowledge base lands, B re-enters at “build” candidacy; if it doesn't, that tells you something too.
- Extraction accuracy & pipeline reality → a fixed-scope review of use case A against the page-3 gate conditions, on your real documents.
- EU-only / no-transfer processing → a Load-Bearing Review (security & EU AI Act), scoped by the exposure tier already assigned above.
- Taking A from verdict to production → implementation and delivery: one workflow, built, integrated and handed over, measured against the page-3 metric.
This document is yours either way.
Appendix · The Gate Zero rubric
The anchors are the method. Nothing behind the curtain.
Every dimension is scored 1–5 against observable anchors — the same anchors in every readout, fixed before scoring begins. Score 5 and score 1 are shown here for each dimension; the full anchor set travels with the readout.
- Value at stake
- 5: euros or hours attached to a named workflow, from the client's own figures. 1: “efficiency” with no workflow named.
- Data reality
- 5: grounding content exists, is clean, and is accessible today. 1: it exists in someone's head.
- Grounding & retrieval
- 5: structured, current, chunkable content a retrieval system can serve at production quality. 1: fragmented or stale content no retrieval tuning can rescue.
- Error tolerance vs stakes
- 5: wrong outputs land on a reviewer with time to catch them. 1: wrong outputs reach a customer, a regulator, or a ledger directly.
- Validation & verifiability
- 5: correctness checkable automatically or by cheap sampling. 1: no way to know it's wrong until it costs something.
- Cost, latency & ownership
- 5: unit economics hold at 2× volume and a named owner exists. 1: per-query cost unknown and nobody would be paged.
- Security & EU AI Act exposure
- 5: no personal data, minimal-risk tier, deployment unconstrained. 1: personal or regulated data with an unresolved transfer, or Annex III adjacency unexamined.
Exposure scoring cross-references the EU AI Act (Article 6 / Annex III), the NIST AI Risk Management Framework, and the OWASP Top 10 for LLM applications. Declared assumptions from the session travel with the readout as a register.