Skip to content
Swift IdeasStart a build
Swift Ideas
Start a build

AI inside your product

AI features that check their own work.

We build features that pull data out of documents and screenshots, answer questions from your own data and draft text with its sources cited, inside products people pay for. The model does the reading; code does the deciding. Fixed rules build the record, a checker rejects what cannot be true, and a stronger model is called only when a cheaper one cannot pass.

What clients arrive with

  • Our extraction works on the demo files and falls over on real uploads.
  • The model invents fields, references or numbers, and users notice.
  • Our AI bill went up and nobody changed any code.
  • We changed the prompt and cannot tell whether it is better or worse.
  • Users do not trust the answer because they cannot see where it came from.

What you get

Extraction with a rules engine
The model reads primitives from a screenshot, PDF or CSV. Deterministic code derives the final record from them, so business rules live in tested functions, not in a prompt.
Validator, repair and escalation
A validator names what is wrong with each record. Fixable issues go back to the cheap model as a targeted repair prompt; complex or contradictory cases escalate to a stronger model with the reason logged. Each rung records its cost, so the ladder is tuned on evidence.
Retrieval and generation with citations
Answers grounded in a curated source set the model is told to cite from and nothing else. SwiftQMS drafts SOPs against 26 verified regulatory references and marks every output as a draft for qualified review.
Evals you can run before every release
Fixture sets built from real inputs, including the ones that fail today, scored on each change. Benchmark scripts compare candidate models on your inputs before anyone switches provider.
Model cost control
Every call records tokens and cost. A pricing check compares hardcoded rates with what the provider charges today and exits non-zero when a rate moves, so a lapsed promotional price cannot raise the bill unnoticed.
Review built into the interface
Low-confidence results are flagged for review, users can report a bad extraction from the screen where they saw it, and each report is kept so it can become a test case.
Data handling fit for your sector
Zero Data Retention API configuration where the provider offers it, UK or EU hosting, and organisation-scoped access enforced in Postgres.

How the work runs

  1. Collect the difficult inputs

    We gather a sample of real uploads and questions, weighted towards the ones that break today. That set becomes the first eval.

  2. Split the model from the rules

    We decide what must be deterministic code and what genuinely needs a model, then write the validator before the prompt.

  3. Build the ladder and score it

    Cheap path first, repair second, escalation last. Each rung is scored on accuracy and cost against the eval set before it ships.

  4. Ship with meters on

    Cost per call, escalation reasons and user correction rate go on a dashboard from day one, so the next change is chosen on evidence.

How engagements start

  • Feasibility spike: your real inputs through a first pipeline, with an accuracy and cost read-out, before any build commitment.
  • Prototype to production in 4–8 weeks for one feature, including evals, cost logging and review UX.
  • Retained iteration: we watch escalation and correction rates with you and tune prompts, rules and providers against them.

Built with

  • Anthropic Claude API
  • Gemini
  • OpenRouter
  • Google Cloud Vision
  • TypeScript
  • Next.js
  • Supabase (Postgres, Storage)
  • Vercel
  • Vitest
  • Playwright

Where each number comes from

Each figure names where it comes from.

verified regulatory references the model may cite
26
Source: SwiftQMS: counted in the reference list our Responsible Person maintains
profile questions, answered once
18
Source: SwiftQMS: the product specification, and counted in the built questionnaire
drafts per organisation per day, refused if usage cannot be checked
50
Source: SwiftQMS: the limit set in the product's code

The work behind it

Questions, answered

Something else on your mind? hello@swiftideas.com

Do you fine-tune models?

Rarely, because it is seldom the cheapest fix. We start with prompting, retrieval and deterministic checks, which are faster to change and easier to test. If your data and volume justify fine-tuning, we will show you the numbers first.

How do you stop the model inventing things?

By not letting it have the final word. Records are derived and validated in code, citations are restricted to a curated source set, and anything that fails validation is repaired, escalated or shown to a person, never silently saved.

How do you measure whether it is working?

An eval set scored on every change, plus three production numbers: how often a record needs escalation, how often a user corrects it, and what each call costs.

Can we change model provider later?

Yes. Providers sit behind one interface, so a switch is a configuration change plus a benchmark run on your own inputs, not a rewrite.

What happens to our users' data?

We configure Zero Data Retention with the model provider where it is offered, host in the UK or EU on request, and enforce organisation-level access in the database rather than in application code alone.

What will it cost to run?

It depends on volume and the mix of cheap and escalated calls. We model cost per call during the spike, log it in production, and set up an alert for when provider pricing changes.

Bring the idea. We’ll build all of it.

Tell us what you want to exist and who it is for. The two people who reply are the two people who will build it.