The Business Diagnostic
Where this sits. This is the Gate Zero verdict — the go/no-go stage that follows the Business Diagnostic, and only when the diagnosis concludes something should be built.Illustrative sample. “Arvela Logistics” is a fictional company and every number below is invented. The method, the scoring, and the format are exactly what you receive — built on your numbers instead.

A sample Gate Zero Readout.

Page 1 · Decision memo

Gate Zero Readout — Arvela Logistics · March 2026 · Prepared by Luca-Ștefan Tamaș, LT Strategy Partners

Ship the shipment-document assistant. Consolidate the knowledge base before any customer chatbot. Don't build the pricing engine.

Of the three use cases assessed, one is ready to build now: an assistant that reads booking emails and attached shipment documents and drafts structured entries for the operations team to confirm. The other two fail Gate Zero today — one temporarily, one on design.

  • The document assistant clears the gate. Five operations staff spend roughly nine hours a week each re-keying data from PDFs and emails — about €54,000 a year at a conservative base case, using Arvela's own loaded-cost figures. The grounding content exists, is accessible today, and every output is confirmed by a human before it enters the TMS — so the error budget is honest.
  • The support chatbot is “not yet”, not “no”. The answers customers need live in seventeen inboxes and a wiki last updated fourteen months ago. Built today, the assistant would confidently repeat stale answers. The unlock is organizational, not technical — and it's named on page 4.
  • The pricing engine fails on stakes, not on ambition.A wrong extraction gets caught by the operator who confirms it. A wrong price goes straight to a customer. With spot-quote history living in email threads, the error rate cannot be validated to the tolerance the workflow demands — this is where AI won't pay off for Arvela as scoped.

What leadership already suspected, confirmed with numbers: the re-keying burden your operations lead has flagged for two years is the strongest AI case in the company — stronger than either of the two ideas that came from vendor conversations.

Owner & 90-day stepOperations director. Scope the document assistant against the gate checklist on page 3; in parallel, assign one owner to the knowledge base with a 90-day consolidation target.

Page 2 · The opportunity portfolio

Three use cases, one map.

Dimension scores at walkthrough depth

Gate Zero dimensionA · DocumentsB · ChatbotC · Pricing
Value at stakeGAA
Data realityGRR
Grounding & retrievalGRR
Error tolerance vs stakesGGR
Validation & verifiabilityAAR
Cost, latency & ownershipGAA
Security & EU AI Act exposureAAG
  • G On footing — no blocker found at walkthrough depth
  • A Addressable gap — named condition must be met
  • R Gating problem — blocks a build verdict as scoped

Map badges show each use case's provisional EU AI Act tier. Additional flags — like margin-critical on C — mark business stakes; they are not AI Act tiers.

  • A / Validation (amber): extraction accuracy needs a 200-document ground-truth set before launch — one week of operator time, scoped in the gate checklist.
  • A / Exposure (amber): shipping documents carry personal data; processing must stay in the EU, and the extraction step should run self-hosted or in an EU region with no third-country transfer.
  • B / Value (amber): the interruption cost is real but unmeasured — the base case needs two weeks of ticket counting before it can be banked.
  • B / Data reality (red): no single, current source of the answers exists — they're scattered across seventeen inboxes and a wiki untouched for fourteen months.
  • B / Grounding & retrieval (red): fragmented, stale content is precisely what retrieval tuning cannot rescue — the score follows the anchor until the knowledge base is consolidated.
  • B / Validation (amber): correctness is checkable by sampling against the ops team — but only once a consolidated source defines what a correct answer is.
  • B / Cost & ownership (amber): query volumes and per-answer economics can't be modeled until the assistant has a source to answer from, and no owner is named yet.
  • B / Exposure (amber): customer-facing, so EU AI Act transparency obligations apply, and transcripts would carry customer personal data — both addressable at design time.
  • C / Value (amber): the margin upside is plausible but can't be validated from unstructured quote history — value that can't be measured can't be banked.
  • C / Data reality (red): spot-quote history lives in email threads; no structured record exists to train or validate against.
  • C / Grounding & retrieval (red): the pricing context that would ground a quote — lane, season, client history — isn't retrievable from any current system.
  • C / Error tolerance (red): a wrong price reaches a customer with no human between the model and the harm — the workflow as designed has no error budget.
  • C / Validation (red): with no ground truth, the error rate cannot even be measured to the tolerance the workflow demands.
  • C / Cost & ownership (amber): per-quote economics look fine on paper, but they're meaningless while the error rate is unmeasurable.

The binding constraint at Arvela: data consolidation, not ambition, gates everything after use case A.

Gate Zero weights are fixed before scoring, by design — what matters most to your business is agreed with you earlier, during the diagnosis. No aggregate score is reported — the lowest dimension gates; averages hide.

Page 3 · Top use case — brief & formal gate

A · The shipment-document assistant.

Business process
Inbound booking requests: email + attached PDFs (bills of lading, packing lists, customs forms) re-keyed into the TMS by the operations team.
Impact, conservative base case
5 operators × ~9 hrs/week re-keying × Arvela's loaded cost ≈ €54,000/year, before error-correction time. Assumption declared: volumes from the ops lead's estimate, not measured logs.
Buy / integrate / build
Integrate: an extraction model behind a review UI, wired to the existing TMS API. No custom model training warranted at this volume.
Data readiness
Green at walkthrough depth — documents accessible in the shared mailbox today; ground-truth set needed pre-launch (see gate).
Cost band & time to production
Low five figures, 8–12 weeks to a production system with validation, monitoring, and handover. Range, not a quote — a fixed price follows a scoped audit.
Success metric
Hours of re-keying per 100 shipments, measured before and ~30 days after launch.

The gate

Build now — subject to the checklist below.

  • Capability target: ≥95% field-level extraction accuracy on the 200-document ground-truth set before go-live; every output human-confirmed until then and sampled after.
  • Safety target: no auto-commit to the TMS — operator confirmation stays in the loop until three consecutive weeks inside the error budget.
  • Cost at scale: unit cost per processed shipment stays under the agreed ceiling at 2× current volume; measured, not assumed.
  • Sign-off domains: operations owner · security/data-protection · engineering · the operators who use it.
  • Walk-away conditions: if the ground-truth set can't reach the accuracy floor, or EU-only processing can't be met at acceptable cost, stop — the readout's verdict flips to “not yet” and says why.
  • What would change this verdict: a TMS migration mid-build, or volumes 3× below the estimate used in the base case.

Confidence: scores marked at walkthrough depth only — verified against described systems, not inspected pipelines. The paid audit tests these against the actual mailbox and TMS before any build.

Page 4 · Where AI won't pay off for Arvela

The honest no's — and what would change them.

Not yet B · Customer support assistant

The value is real — status queries interrupt the same team that re-keys documents. But the grounding content fails Gate Zero: the answers live in seventeen inboxes and a wiki last touched fourteen months ago. An assistant built on that would answer confidently and wrongly, in front of customers, under the EU AI Act's transparency obligations for customer-facing bots.The unlock: one owner, one consolidated knowledge base, a 90-day target — then this use case re-enters the gate with a realistic path to “build”.

Don't build C · Dynamic pricing engine

Not because the model can't produce prices — because the workflow has no error budget and no validation path. A wrong extraction is caught by an operator; a wrong price reaches a customer directly and costs margin or trust. The spot-quote history that would train and validate it lives in unstructured email threads, so the error rate can't even be measured to the tolerance required.The cheaper answer that survives scrutiny: a rules-based margin floor with quarterly review — nearly free, no surprise behavior. What would change the verdict:two years of structured quote/outcome data and a human approval step that doesn't erase the speed advantage.

Page 5 · Gaps, next steps, and the map forward

What happens next — with us or without us.

  1. Assign the knowledge-base owner (yours to do, this week). No engagement needed — one name, one 90-day consolidation target, and use case B stops being blocked.
  2. Build the ground-truth set (yours to do, ~1 week of operator time). 200 recent shipments, fields verified by hand. Whoever builds the assistant will need it — us or anyone else.
  3. Scope the document assistant against the page-3 gate. The open questions — actual extraction accuracy on your documents, EU-only processing cost — are exactly what a fixed-scope review of the candidate against these gate conditions tests, on your real documents.
  4. Re-gate the chatbot in 90 days. If the knowledge base lands, B re-enters at “build” candidacy; if it doesn't, that tells you something too.

If you want help, this is how each gap maps

  • Extraction accuracy & pipeline reality → a fixed-scope review of use case A against the page-3 gate conditions, on your real documents.
  • EU-only / no-transfer processing → a Load-Bearing Review (security & EU AI Act), scoped by the exposure tier already assigned above.
  • Taking A from verdict to production → implementation and delivery: one workflow, built, integrated and handed over, measured against the page-3 metric.

This document is yours either way.

Appendix · The Gate Zero rubric

The anchors are the method. Nothing behind the curtain.

Every dimension is scored 1–5 against observable anchors — the same anchors in every readout, fixed before scoring begins. Score 5 and score 1 are shown here for each dimension; the full anchor set travels with the readout.

Value at stake
5: euros or hours attached to a named workflow, from the client's own figures. 1: “efficiency” with no workflow named.
Data reality
5: grounding content exists, is clean, and is accessible today. 1: it exists in someone's head.
Grounding & retrieval
5: structured, current, chunkable content a retrieval system can serve at production quality. 1: fragmented or stale content no retrieval tuning can rescue.
Error tolerance vs stakes
5: wrong outputs land on a reviewer with time to catch them. 1: wrong outputs reach a customer, a regulator, or a ledger directly.
Validation & verifiability
5: correctness checkable automatically or by cheap sampling. 1: no way to know it's wrong until it costs something.
Cost, latency & ownership
5: unit economics hold at 2× volume and a named owner exists. 1: per-query cost unknown and nobody would be paged.
Security & EU AI Act exposure
5: no personal data, minimal-risk tier, deployment unconstrained. 1: personal or regulated data with an unresolved transfer, or Annex III adjacency unexamined.

Exposure scoring cross-references the EU AI Act (Article 6 / Annex III), the NIST AI Risk Management Framework, and the OWASP Top 10 for LLM applications. Declared assumptions from the session travel with the readout as a register.

This is the document your leadership team receives — built on your numbers, your systems, and your use cases.

Start with the Business Diagnostic