AI inside your product
AI features that check their own work.
We build features that pull data out of documents and screenshots, answer questions from your own data and draft text with its sources cited, inside products people pay for. The model does the reading; code does the deciding. Fixed rules build the record, a checker rejects what cannot be true, and a stronger model is called only when a cheaper one cannot pass.
What clients arrive with
- Our extraction works on the demo files and falls over on real uploads.
- The model invents fields, references or numbers, and users notice.
- Our AI bill went up and nobody changed any code.
- We changed the prompt and cannot tell whether it is better or worse.
- Users do not trust the answer because they cannot see where it came from.
What you get
- Extraction with a rules engine
- The model reads primitives from a screenshot, PDF or CSV. Deterministic code derives the final record from them, so business rules live in tested functions, not in a prompt.
- Validator, repair and escalation
- A validator names what is wrong with each record. Fixable issues go back to the cheap model as a targeted repair prompt; complex or contradictory cases escalate to a stronger model with the reason logged. Each rung records its cost, so the ladder is tuned on evidence.
- Retrieval and generation with citations
- Answers grounded in a curated source set the model is told to cite from and nothing else. SwiftQMS drafts SOPs against 26 verified regulatory references and marks every output as a draft for qualified review.
- Evals you can run before every release
- Fixture sets built from real inputs, including the ones that fail today, scored on each change. Benchmark scripts compare candidate models on your inputs before anyone switches provider.
- Model cost control
- Every call records tokens and cost. A pricing check compares hardcoded rates with what the provider charges today and exits non-zero when a rate moves, so a lapsed promotional price cannot raise the bill unnoticed.
- Review built into the interface
- Low-confidence results are flagged for review, users can report a bad extraction from the screen where they saw it, and each report is kept so it can become a test case.
- Data handling fit for your sector
- Zero Data Retention API configuration where the provider offers it, UK or EU hosting, and organisation-scoped access enforced in Postgres.
How the work runs
Collect the difficult inputs
We gather a sample of real uploads and questions, weighted towards the ones that break today. That set becomes the first eval.
Split the model from the rules
We decide what must be deterministic code and what genuinely needs a model, then write the validator before the prompt.
Build the ladder and score it
Cheap path first, repair second, escalation last. Each rung is scored on accuracy and cost against the eval set before it ships.
Ship with meters on
Cost per call, escalation reasons and user correction rate go on a dashboard from day one, so the next change is chosen on evidence.
How engagements start
- Feasibility spike: your real inputs through a first pipeline, with an accuracy and cost read-out, before any build commitment.
- Prototype to production in 4–8 weeks for one feature, including evals, cost logging and review UX.
- Retained iteration: we watch escalation and correction rates with you and tune prompts, rules and providers against them.
Built with
Where each number comes from
Each figure names where it comes from.
- verified regulatory references the model may cite
- 26
- Source: SwiftQMS: counted in the reference list our Responsible Person maintains
- profile questions, answered once
- 18
- Source: SwiftQMS: the product specification, and counted in the built questionnaire
- drafts per organisation per day, refused if usage cannot be checked
- 50
- Source: SwiftQMS: the limit set in the product's code
The work behind it
Questions, answered
Something else on your mind? hello@swiftideas.com
Do you fine-tune models?
Rarely, because it is seldom the cheapest fix. We start with prompting, retrieval and deterministic checks, which are faster to change and easier to test. If your data and volume justify fine-tuning, we will show you the numbers first.
How do you stop the model inventing things?
By not letting it have the final word. Records are derived and validated in code, citations are restricted to a curated source set, and anything that fails validation is repaired, escalated or shown to a person, never silently saved.
How do you measure whether it is working?
An eval set scored on every change, plus three production numbers: how often a record needs escalation, how often a user corrects it, and what each call costs.
Can we change model provider later?
Yes. Providers sit behind one interface, so a switch is a configuration change plus a benchmark run on your own inputs, not a rewrite.
What happens to our users' data?
We configure Zero Data Retention with the model provider where it is offered, host in the UK or EU on request, and enforce organisation-level access in the database rather than in application code alone.
What will it cost to run?
It depends on volume and the mix of cheap and escalated calls. We model cost per call during the spike, log it in production, and set up an alert for when provider pricing changes.

