One person, working to a written scope. Every job has a fixed price, a stated turnaround, an exceptions file and a note of the rules applied, so the result is reproducible without me.
Your messy export, cleaned once and properly — $149
You have an export that nobody trusts: dates in three formats, amounts with currency symbols and thousands separators, duplicate rows, blank lines, and a column someone renamed last quarter. Send it, and you get back one clean file against the schema you confirm at the start, plus a separate file of every row we would not guess at, each with the reason.
What is included
- up to 5 files per job and up to 50,000 rows in total, CSV or XLSX
- one target schema, agreed before we start
- header normalisation, type normalisation (dates to ISO, money to numbers), whitespace and case cleanup
- deduplication on one key you choose
- one re-run included after you review the result
What is not included
- scanned files, photographs of tables and PDFs — this offer reads text, not images
- more than one target schema per job (a second schema is a second job)
- inventing values: a row we cannot map with certainty goes to exceptions, always
- deciding what your data means: we normalise the shape, you own the rules
- anything requiring access to your systems — you send files, we send files
What I need from you
- the file or files, as CSV or XLSX
- the target schema: the column names you want out, and the type of each
- the deduplication key (the column that makes a row unique)
- the date order your system writes, if you know it — we infer it from the file and the file wins, but knowing saves exceptions
- the currency of any amounts column
What you get
- clean.csv
- clean.xlsx
- exceptions.csv, one row per refusal with the reason in plain words
- a one-page note of every rule applied, so the result is reproducible without us
Turnaround: 2 business days from the moment the schema is confirmed
How it is checked before it is sent
- the file's own date order is inferred from unambiguous values before a single row is parsed; a file that proves both orders has its ambiguous dates refused, not guessed
- every number is re-checked against the source token: same digits, same sign, same decimal places
- row counts are reconciled: rows in = rows delivered + exceptions + duplicates removed, with no unexplained loss
- the exceptions file is read end to end by a person before anything is sent
- the deduplication key is checked for collisions that are not true duplicates before any row is dropped
When something cannot be read
- a row we cannot map with certainty is never guessed at: it goes to exceptions with the reason
- measured on real published spend files, exceptions run at about 0.5% of rows, and every one of them was a genuinely unreadable value
- typical causes, all real: a number exported as ########, a currency symbol corrupted before we ever saw it, a date that could be day-first or month-first in a file that proves neither
If I get it wrong
if the delivered file does not match the schema you confirmed, we redo the job at our cost, once, within one business day. If the second attempt still does not match the confirmed schema, the fee is refunded in full.
What this rests on
- REAL_WORLD_BENCHMARK 2026-09-20: 4,464 held-out rows from 8 published UK government spend files (Open Government Licence v3.0, 4 publishers) — 100% field accuracy, 99.51% straight-through, 0.49% exceptions, 0 silent errors
- SYNTHETIC_BENCHMARK: 100% straight-through and 0 undetected errors on single-format files; 81.2% straight-through and 28.7% exceptions on deliberately mixed-format files
$149 per job — up to 5 files, up to 50,000 rows, one schema, one re-run included
Not for you if: your data is in scans or photographs (we would be guessing, and we do not guess); you want the rules decided for you — you own what the data means; you need it inside a day: the turnaround is 2 business days and we would rather say so.
The same report, every period, without you rebuilding it — $99 a month
Someone on your team rebuilds the same report from the same export every month, and it takes half a day and occasionally goes wrong. Drop the file in, get the standardised report back the same day, every period, with the same rules applied every time.
What is included
- one recurring source file in a fixed format, CSV or XLSX, up to 100,000 rows per period
- one agreed report layout: the same columns, the same groupings, the same totals, every period
- the exceptions file each period, so you can see what changed in your source data
- the rules note, updated whenever a rule changes
What is not included
- a source format change is a new setup, priced separately — we will tell you the moment we see one
- more than one source file per period
- analysis, commentary or interpretation: this is the same report, reliably, not a new one each month
- scanned or image inputs
What I need from you
- one sample export in the exact format you will send each period
- the report layout you want: columns, groupings, totals
- the day of the period you will send it
- who receives the output
What you get
- report.xlsx in the agreed layout
- exceptions.csv for the period
- a short note of anything that changed in your source data since last period
Turnaround: same business day, every period, provided the file arrives in the agreed format
How it is checked before it is sent
- the file's own date order is inferred from unambiguous values before a single row is parsed; a file that proves both orders has its ambiguous dates refused, not guessed
- every number is re-checked against the source token: same digits, same sign, same decimal places
- row counts are reconciled: rows in = rows delivered + exceptions + duplicates removed, with no unexplained loss
- the exceptions file is read end to end by a person before anything is sent
- this period's totals are compared with last period's, and a swing beyond the agreed band is read by a person before the report goes out
When something cannot be read
- rows that cannot be mapped go to the exceptions file, never into the totals
- a change in your source format stops the run and produces a message, rather than a report built on a guess
If I get it wrong
a report that does not match the agreed layout is corrected the same day, at our cost. A period we cannot deliver at all is not billed for that period.
What this rests on
- REAL_WORLD_BENCHMARK 2026-09-20: the same pipeline on 4,464 held-out rows of published UK government spend data — 100% field accuracy, 0 silent errors
- setup pays back in 1 month at $99; delivery costs $22.74 per period against a $99 price
$99 per month — one recurring report, same day, every period
Not for you if: your export format changes often — each change is a new setup; you need the analysis, not the assembly; one period's delay would be a serious problem: we are one person and we say so up front.
100 supplier invoices, typed into one spreadsheet — $149
You receive invoices as PDFs and someone types them into a spreadsheet. Send the batch, get back one validated spreadsheet with the fields you agreed, every total cross-checked against its own line items, and a separate file of the documents we would not guess at.
What is included
- up to 100 digital PDFs per batch — files with a real text layer, as produced by accounting software
- the agreed fields per document: invoice number, date, currency, subtotal, tax, total, line count
- an arithmetic cross-check on every document: line items and tax must reconcile to the stated total
- one re-run included
What is not included
- scanned or photographed invoices — those need OCR, which this offer does not include and which we have not measured
- handwritten documents
- line-item detail beyond the count (a per-line extract is a different job)
- posting anything into your accounting system
What I need from you
- the PDFs, in one archive
- the field list you want out
- the currency, if your suppliers bill in more than one
- one sample of each invoice layout you expect, if you know them
What you get
- invoices.xlsx with one row per document and the agreed fields
- exceptions.csv listing every document we refused, with the reason
- a note of the layouts seen, so the next batch is faster
Turnaround: 2 business days from receipt of the batch
How it is checked before it is sent
- each document's own arithmetic is checked: subtotal plus tax must equal the stated total, or the document becomes an exception
- every field carries how it was found — a primary label, a secondary label, or computed — and anything below a primary label is read by a person
- the file's own date order is inferred from unambiguous values before a single row is parsed; a file that proves both orders has its ambiguous dates refused, not guessed
- every number is re-checked against the source token: same digits, same sign, same decimal places
- a sample of delivered rows is compared against the source PDFs by eye before the batch is sent
When something cannot be read
- a document whose total does not reconcile is never delivered as if it did: it goes to exceptions with the arithmetic shown
- measured on synthetic batches, 91.3% of 1,000 documents went straight through and every single error was caught by the arithmetic check rather than delivered
- a scanned document in a digital batch is returned as an exception, not run through an OCR path we have not measured
If I get it wrong
any document we got wrong is re-extracted at our cost, and if more than 5% of a batch is wrong we redo the batch. If the second attempt still fails the agreed field list, the fee is refunded in full.
What this rests on
- SYNTHETIC_BENCHMARK: 94% straight-through on 100 documents, 91.3% on 1,000, 72% on deliberately mixed layouts — 100% field accuracy and 0 undetected errors in every run
- NOT YET MEASURED on real invoices: no corpus of real digital invoices with field-level ground truth and a commercial licence has been found (docs/OFFER_COSTING.md §8)
$149 per batch of up to 100 digital PDFs
Not for you if: your invoices are scans or photographs — say so and we will tell you honestly that this offer does not cover it; you need them posted into your ledger; you need per-line detail rather than document totals.
How a job runs
- You send the files and the target schema. If anything is ambiguous, I ask once, before starting.
- The clock starts when the scope is confirmed — not when the files arrive.
- Automated processing, then a human read. The parsing is deterministic; the judgement is not.
- You get the clean file, the exceptions file and the rules note within the stated turnaround.
- One re-run is included. If the second attempt still does not match the confirmed scope, you pay nothing.
Honest limits
- This is one person. There is no team, no queue and no overnight turnaround, and I will say so rather than miss a date.
- Scanned or photographed documents are out of scope. That work needs a different pipeline, which I have not measured, so I do not sell it.
- There are no testimonials on this page because there are no customers yet. The quality claims point at measurements, not opinions.
Get in touch
Email pradomation@gmail.com with what you have and what you need out of it. Include a sample if you can — five rows is enough to tell you whether it is in scope.
No newsletter, no calls booked automatically, no follow-up sequence. You write, I answer.
Privacy
This site sets no cookies, loads nothing from anywhere else, and uses no analytics service. It records four
things, on its own servers, and nothing else: the src label in the link you arrived through (if there
was one), which offer sections you scrolled to, whether you clicked the enquiry link, and when. No IP address, no
browser or device details, no identifier, no session, nothing that could be linked back to you or to another visit.
If you send files for a job: they are used only to do that job, are not used to train anything, are not shared with anyone, and are deleted on request or within 30 days of delivery, whichever comes first. Delivered results and an exceptions file are kept only until you confirm receipt.