← The Business Diagnostic
Where this sits. This is the Gate Zero verdict — the
go/no-go stage that follows the Business Diagnostic, and only when the
diagnosis concludes something should be built.
Illustrative sample. “Arvela Logistics” is a fictional
company and every number below is invented. The method, the scoring,
and the format are exactly what you receive — built on your numbers
instead.
A sample Gate Zero Readout.
Page 1 · Decision memo
Gate Zero Readout — Arvela Logistics · March 2026 · Prepared by
Luca-Ștefan Tamaș, LT Strategy Partners
Ship the shipment-document assistant. Consolidate the knowledge base
before any customer chatbot. Don't build the pricing engine.
Of the three use cases assessed, one is ready to build now: an
assistant that reads booking emails and attached shipment documents
and drafts structured entries for the operations team to confirm.
The other two fail Gate Zero today — one temporarily, one on design.
- The document assistant clears the gate. Five
operations staff spend roughly nine hours a week each re-keying
data from PDFs and emails — about €54,000 a year at a conservative
base case, using Arvela's own loaded-cost figures. The grounding
content exists, is accessible today, and every output is confirmed
by a human before it enters the TMS — so the error budget is honest.
- The support chatbot is “not yet”, not “no”. The
answers customers need live in seventeen inboxes and a wiki last
updated fourteen months ago. Built today, the assistant would
confidently repeat stale answers. The unlock is organizational,
not technical — and it's named on page 4.
- The pricing engine fails on stakes, not on ambition.
A wrong extraction gets caught by the operator who confirms it. A
wrong price goes straight to a customer. With spot-quote history
living in email threads, the error rate cannot be validated to the
tolerance the workflow demands — this is where AI won't pay off
for Arvela as scoped.
What leadership already suspected, confirmed with numbers: the
re-keying burden your operations lead has flagged for two years is
the strongest AI case in the company — stronger than either of the
two ideas that came from vendor conversations.
Owner & 90-day step
Operations director. Scope the document assistant against the gate
checklist on page 3; in parallel, assign one owner to the knowledge
base with a 90-day consolidation target.
Page 2 · The opportunity portfolio
Three use cases, one map.
Value at stake → Production viability → Not yet — close the gap first
Build now
Don't build — and here's why
Quick wins, modest stakes
A · Document assistant AI Act: minimal
B · Support chatbot AI Act: limited
C · Pricing engine AI Act: minimal margin-critical
Dimension scores at walkthrough depth
- G On footing — no blocker found at walkthrough depth
- A Addressable gap — named condition must be met
- R Gating problem — blocks a build verdict as scoped
Map badges show each use case's provisional EU AI Act tier. Additional
flags — like margin-critical on C — mark business stakes; they
are not AI Act tiers.
- A / Validation (amber): extraction accuracy needs a 200-document ground-truth set before launch — one week of operator time, scoped in the gate checklist.
- A / Exposure (amber): shipping documents carry personal data; processing must stay in the EU, and the extraction step should run self-hosted or in an EU region with no third-country transfer.
- B / Value (amber): the interruption cost is real but unmeasured — the base case needs two weeks of ticket counting before it can be banked.
- B / Data reality (red): no single, current source of the answers exists — they're scattered across seventeen inboxes and a wiki untouched for fourteen months.
- B / Grounding & retrieval (red): fragmented, stale content is precisely what retrieval tuning cannot rescue — the score follows the anchor until the knowledge base is consolidated.
- B / Validation (amber): correctness is checkable by sampling against the ops team — but only once a consolidated source defines what a correct answer is.
- B / Cost & ownership (amber): query volumes and per-answer economics can't be modeled until the assistant has a source to answer from, and no owner is named yet.
- B / Exposure (amber): customer-facing, so EU AI Act transparency obligations apply, and transcripts would carry customer personal data — both addressable at design time.
- C / Value (amber): the margin upside is plausible but can't be validated from unstructured quote history — value that can't be measured can't be banked.
- C / Data reality (red): spot-quote history lives in email threads; no structured record exists to train or validate against.
- C / Grounding & retrieval (red): the pricing context that would ground a quote — lane, season, client history — isn't retrievable from any current system.
- C / Error tolerance (red): a wrong price reaches a customer with no human between the model and the harm — the workflow as designed has no error budget.
- C / Validation (red): with no ground truth, the error rate cannot even be measured to the tolerance the workflow demands.
- C / Cost & ownership (amber): per-quote economics look fine on paper, but they're meaningless while the error rate is unmeasurable.
The binding constraint at Arvela: data consolidation, not
ambition, gates everything after use case A.
Page 3 · Top use case — brief & formal gate
A · The shipment-document assistant.
- Business process
- Inbound booking requests: email + attached PDFs (bills of lading, packing lists, customs forms) re-keyed into the TMS by the operations team.
- Impact, conservative base case
- 5 operators × ~9 hrs/week re-keying × Arvela's loaded cost ≈ €54,000/year, before error-correction time. Assumption declared: volumes from the ops lead's estimate, not measured logs.
- Buy / integrate / build
- Integrate: an extraction model behind a review UI, wired to the existing TMS API. No custom model training warranted at this volume.
- Data readiness
- Green at walkthrough depth — documents accessible in the shared mailbox today; ground-truth set needed pre-launch (see gate).
- Cost band & time to production
- Low five figures, 8–12 weeks to a production system with validation, monitoring, and handover. Range, not a quote — a fixed price follows a scoped audit.
- Success metric
- Hours of re-keying per 100 shipments, measured before and ~30 days after launch.
The gate
Build now — subject to the checklist below.
- Capability target: ≥95% field-level extraction accuracy on the 200-document ground-truth set before go-live; every output human-confirmed until then and sampled after.
- Safety target: no auto-commit to the TMS — operator confirmation stays in the loop until three consecutive weeks inside the error budget.
- Cost at scale: unit cost per processed shipment stays under the agreed ceiling at 2× current volume; measured, not assumed.
- Sign-off domains: operations owner · security/data-protection · engineering · the operators who use it.
- Walk-away conditions: if the ground-truth set can't reach the accuracy floor, or EU-only processing can't be met at acceptable cost, stop — the readout's verdict flips to “not yet” and says why.
- What would change this verdict: a TMS migration mid-build, or volumes 3× below the estimate used in the base case.
Confidence: scores marked at walkthrough depth only — verified
against described systems, not inspected pipelines. The paid audit
tests these against the actual mailbox and TMS before any build.
Page 4 · Where AI won't pay off for Arvela
The honest no's — and what would change them.
Not yet B · Customer support assistant
The value is real — status queries interrupt the same team that
re-keys documents. But the grounding content fails Gate Zero: the
answers live in seventeen inboxes and a wiki last touched fourteen
months ago. An assistant built on that would answer confidently
and wrongly, in front of customers, under the EU AI Act's
transparency obligations for customer-facing bots.
The unlock: one owner, one consolidated knowledge
base, a 90-day target — then this use case re-enters the gate with
a realistic path to “build”.
Don't build C · Dynamic pricing engine
Not because the model can't produce prices — because the workflow
has no error budget and no validation path. A wrong extraction is
caught by an operator; a wrong price reaches a customer directly
and costs margin or trust. The spot-quote history that would train
and validate it lives in unstructured email threads, so the error
rate can't even be measured to the tolerance required.
The cheaper answer that survives scrutiny: a
rules-based margin floor with quarterly review — nearly free, no
surprise behavior. What would change the verdict:
two years of structured quote/outcome data and a human approval
step that doesn't erase the speed advantage.
Page 5 · Gaps, next steps, and the map forward
What happens next — with us or without us.
- Assign the knowledge-base owner (yours to do, this week). No engagement needed — one name, one 90-day consolidation target, and use case B stops being blocked.
- Build the ground-truth set (yours to do, ~1 week of operator time). 200 recent shipments, fields verified by hand. Whoever builds the assistant will need it — us or anyone else.
- Scope the document assistant against the page-3 gate. The open questions — actual extraction accuracy on your documents, EU-only processing cost — are exactly what a fixed-scope review of the candidate against these gate conditions tests, on your real documents.
- Re-gate the chatbot in 90 days. If the knowledge base lands, B re-enters at “build” candidacy; if it doesn't, that tells you something too.
- Extraction accuracy & pipeline reality → a fixed-scope review of use case A against the page-3 gate conditions, on your real documents.
- EU-only / no-transfer processing → a Load-Bearing Review (security & EU AI Act), scoped by the exposure tier already assigned above.
- Taking A from verdict to production → implementation and delivery: one workflow, built, integrated and handed over, measured against the page-3 metric.
This document is yours either way.
Appendix · The Gate Zero rubric
The anchors are the method. Nothing behind the curtain.
Every dimension is scored 1–5 against observable anchors — the same
anchors in every readout, fixed before scoring begins. Score 5 and
score 1 are shown here for each dimension; the full anchor set
travels with the readout.
- Value at stake
- 5: euros or hours attached to a named workflow, from the client's own figures. 1: “efficiency” with no workflow named.
- Data reality
- 5: grounding content exists, is clean, and is accessible today. 1: it exists in someone's head.
- Grounding & retrieval
- 5: structured, current, chunkable content a retrieval system can serve at production quality. 1: fragmented or stale content no retrieval tuning can rescue.
- Error tolerance vs stakes
- 5: wrong outputs land on a reviewer with time to catch them. 1: wrong outputs reach a customer, a regulator, or a ledger directly.
- Validation & verifiability
- 5: correctness checkable automatically or by cheap sampling. 1: no way to know it's wrong until it costs something.
- Cost, latency & ownership
- 5: unit economics hold at 2× volume and a named owner exists. 1: per-query cost unknown and nobody would be paged.
- Security & EU AI Act exposure
- 5: no personal data, minimal-risk tier, deployment unconstrained. 1: personal or regulated data with an unresolved transfer, or Annex III adjacency unexamined.
Exposure scoring cross-references the EU AI Act (Article 6 / Annex
III), the NIST AI Risk Management Framework, and the OWASP Top 10
for LLM applications. Declared assumptions from the session travel
with the readout as a register.