A deterministic core, a model only at the edges
Pricing, policy, classification and reconciliation are code. The model parses language and explains results. It never decides something that has a right answer.
BUILT, NOT PITCHED
Ten problems a business already pays people to solve, built as working software rather than slideware. Tax reconciliation, agent governance, EU AI Act evidence, agentic checkout, nonprofit marketing, retail stock and cost, the support inbox, hiring, sales and shipping. Every number on this page is produced by the system it describes.
01
Social media assistant
THE PROBLEM
Nonprofit marketing is one person, ten channels and no research time. Generic AI copy fails here for one reason: it is unsourced, so nothing it claims can be published without checking it first.
WHAT IT DOES
It researches the organisation before it writes a word, drafts per platform from an editable rule file, and remembers corrections. Tell it once that you are warm and grassroots rather than corporate, and every later draft holds that, across sessions and restarts.
Sources on every claim, so nothing publishable is unattributed. A new platform is a text file, not a deployment.
WHO WOULD BUY IT: agencies and nonprofit networks, per organisation, per month
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
02
myDATA reconciliation
THE PROBLEM
Every Greek business must match its books to the tax office by hand. When the two disagree, somebody finds it document by document, usually late, and the penalty lands on the client.
WHAT IT DOES
It pairs documents on identity, then on a tolerant window so a one-day or one-cent difference is still the same invoice, classifies ten discrepancy types by rule, and explains each one in Greek. It proposes the correction and never files it: a dry run until a person approves.
Ten discrepancy types, classified by rule rather than by a model, so the same period gives the identical answer every run.
WHO WOULD BUY IT: Greek accounting firms, per client company, per month
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
03
Agent control plane
THE PROBLEM
The moment an agent can call tools it can delete, spend, email and leak. Most teams ship that with a prompt asking it to behave. What is missing is the layer that decides, before the call runs, whether it is allowed, and leaves a record either way.
WHAT IT DOES
Allow, deny or route to a human, decided first. Budgets checked before the upstream, so a refused call cannot exhaust a tenant. Injection scanning in both directions. Everything traced, refusals included, because a gateway that logs only what it allowed cannot answer an auditor.
25 of 25 attacks caught with 0 of 25 false positives, and an unmatched tool denied by default so adding a tool never silently adds a capability.
WHO WOULD BUY IT: any company putting agents near production, per seat or per agent
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
04
EU AI Act conformity
THE PROBLEM
The Act is in force and being confidently wrong is the liability. A model asked to write an evidence file from memory produces a fluent document citing articles that do not say what it claims, which is worse than no document at all.
WHAT IT DOES
The model never recalls an article, it resolves one from a pinned, versioned corpus, and a citation that fails to resolve raises rather than finding the nearest match. Closed lists are settled in code at zero cost. A claim survives only if a separate verifier agrees its evidence supports it.
Zero unsourced claims as a structural gate rather than a target, and an edit made after issue is caught by the content fingerprint.
WHO WOULD BUY IT: compliance consultancies and EU deployers, per assessment, then a retainer
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
05
Agent commerce
THE PROBLEM
Agentic checkout is arriving through three incompatible protocols and none has won. A merchant that picks one bets the integration on an unsettled outcome, and one that lets an autonomous buyer negotiate without structural limits gets exactly the sale it deserves.
WHAT IT DOES
One interface behind all three protocols. The margin floor is arithmetic in integer minor units, so no discount at any customer class can price below it. Every mandate is verified before money moves, and nonces, holds and settlements outlive the process.
Oversell, double charge and replay all refused, including across a restart, with no model call anywhere on the money path.
WHO WOULD BUY IT: merchants and platforms selling to autonomous buyers, per transaction or licensed
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
06
Retail ERP suite
THE PROBLEM
A retailer's real numbers live in six tools that disagree. Nobody can answer what a unit actually cost or what the month actually made, because no single system holds both halves.
WHAT IT DOES
Freight and duty are allocated across a shipment by value, so price is set on landed cost rather than on the invoice. Stock is never stored, it is derived from an append-only movement ledger, so every unit traces back to the reception, sale or transfer that produced it.
On one shipment the landed cost ran 8.9% above the supplier price, which is exactly the margin a retailer loses by pricing off the invoice.
WHO WOULD BUY IT: independent retailers and small chains, per location, per month
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
07
Customer support agent
THE PROBLEM
A retail inbox is the same forty questions, and the customer whose parcel is genuinely lost waits behind them. A model that answers the inbox will state a delivery date, fluently and wrongly, which is worse than the backlog.
WHAT IT DOES
Every message is scanned in memory for instructions aimed at the model, then anonymised so that only the redacted text is stored, then sorted by rule into nine intents and answered from the order row or the policy page where the answer is a fact. A reply about an order is rendered from the row and verified against it; a reference the database does not hold escalates to a person. There is no send function anywhere in the package.
156 of 156 synthetic messages routed as planted, 10 of 10 injection attempts caught with 8 of 8 benign lookalikes passed, and 0 replies sent by the agent, because the package has no send path.
WHO WOULD BUY IT: retailers and e-shops with one to five people on support, per inbox, per month
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
08
HR screening assistant
THE PROBLEM
Two hundred CVs for one role, read over a weekend. Every screener sold to fix this either ranks by a model nobody can question or by keywords and calls it AI, and none says out loud that CV screening is Annex III of the EU AI Act.
WHAT IT DOES
Birth date, marital status, children, military service, nationality, photo and name are stripped in code before any scoring reads the document. Six deterministic criteria are scored with the span of the CV that produced each verdict, or refused as no evidence. The decisions table refuses the agent as decider by trigger. Classified high risk by the conformity agent, by rule, and built for it.
360 of 360 verdicts on 60 synthetic CVs as planted, 318 of 318 protected attributes removed, and 30 of 30 matched pairs, identical qualifications with different protected attributes, given identical outcomes.
WHO WOULD BUY IT: companies hiring at volume without a screening system, per role, per month, with the conformity pack attached
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
09
Sales assistant
THE PROBLEM
The proposal is the last one with the name changed, and it still promises delivery in two days because the last one did. A model that writes it invents the delivery date, and a proposal that states a term the company does not offer is a contract it did not mean to sign.
WHAT IT DOES
Prices come from the price list through a floor in integer arithmetic that no discount can cross, terms from the terms table, references from won proposals in the same sector, and every claim is re-read against its row before the draft exists. Follow-ups are a date: last activity plus the stage SLA. The brief cites its sources, and the schema has no LinkedIn kind to cite.
24 of 24 planted unsourced claims stripped before the draft was stored, 7 of 7 floor-breaching lines refused with 0 stored below floor, and 49 of 49 overdue follow-ups found with 0 false alarms.
WHO WOULD BUY IT: small B2B sales teams without an SDR function, per seat, per month
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
10
Logistics
THE PROBLEM
Three carriers, three portals, three vocabularies for "we tried and nobody answered", and a customer asking where the parcel is. The customs form is typed from the delivery note, which was typed from the order, and the mismatch is found at the border.
WHAT IT DOES
A parser per feed format and a code map per carrier put every event on one timeline; a code the map does not know stays unknown and raises for a person, never mapped to the nearest state. Exceptions are rules over dates and codes. Delivery notes, proofs of delivery and customs declarations are rendered from the order rows, re-parsed and diffed, and the documents table refuses anything but a zero-mismatch document.
264 of 264 carrier events normalised from three formats with 9 of 9 unknown codes kept and 0 guessed, 41 of 41 shipments raising exactly the planted exceptions, and 130 of 130 documents re-parsing to their rows with 0 mismatches.
WHO WOULD BUY IT: retailers and small distributors shipping with two or more carriers, per location, per month
Open the dashboard Static build from the running code, synthetic data, opens in a new tab
HOW THEY ARE BUILT
The hard part is not the model. It is everything around it: what the model is allowed to decide, what holds when it is wrong, and whether anyone can check afterwards.
Pricing, policy, classification and reconciliation are code. The model parses language and explains results. It never decides something that has a right answer.
The margin floor, the replay nonce, the tenant scope and the durable ledgers are code that a prompt cannot talk its way around. They hold when detection fails.
Every figure on these screens is produced by the system's own code and re-checked by its test suite. The suites redden under mutation, so green is evidence rather than decoration.
Capabilities are files, not deployments. A non-engineer adds a platform, a skill or a research source and it is live on the next request.
THE PRESENTATION
The same ten systems as a presentation: the problem each one solves, who pays to solve it today, how it works in four moves, and screens of it running. Every figure in it is produced by the system it describes.
Download the presentation PDF, 35 pages, 5.1 MB
Each of these can go to a first pilot as it stands: the interfaces are real, the rules are enforced in code, and every figure comes from the running system. The next step for any of them is a pilot against live data.
Talk about a pilot