Historical records transcription
Scanned or photographed registers, grave record sheets, and other archival records, turned into structured data that can be searched.
I transcribe historical records for local heritage groups, parish archives, museums, and researchers. Most of these collections can be read by a person, yet none of them can be searched. Send a few representative pages and I will read them, then come back with a scope and a quote. I do the work myself, with TextHarvester as the supporting tool, so there is no application to learn and no subscription to manage.
To start, email a few representative pages, including any that are difficult to read, along with the approximate size of the collection and the output you need. I will assess them and quote. The assessment does not include transcribing the collection.
What you receive
One row per agreed record, in a spreadsheet, with one column per agreed field. The columns are settled before transcription starts, since a burial register needs different fields from a set of grave record sheets.
A burial register run typically produces the fields below. Your own column list is agreed during the sample assessment.
entry_id: a stable reference for the row, such asvol2_p017_r004name_raw,age_raw,abode_raw,burial_date_raw,officiant_raw: the record as written, each in a_rawcolumn that keeps the original spellingmarginalia_raw: notes written in the margin, with a pipe marking a line breakpage_numberandrow_index_on_page: where on the page the entry sits, for burial registersfile_name: the source image every row came fromuncertainty_flagsandconfidence_scores: what the transcription was unsure about, and how unsure it was
Every row keeps its source file_name, so any record can be traced back to the image it came from. Burial register rows also carry a page and row reference. If your record type has no page-level reference, the scope will say so.
Send scans or photographs as JPEG images or PDF files. CSV is the standard output and opens directly in Excel, and JSON is available where a system needs it. I do not produce native Excel workbooks (.xlsx). You can search the dataset, import it elsewhere, and check it against the source. Turning it into a searchable website or a map is separate work that I can quote for on its own.
Example project
Real records belong to the organisations that hold them, so a collection can be described publicly only once its delivered scope and publication rights are agreed. Until the scope and rights are settled, the illustration below stands in. The register line in it is invented, and so is every value.
| As it appears on the page | Field | Delivered value |
|---|---|---|
| No. 412. J---n C-----y, abode B------k, buried 14 March 18--, aged 6-, officiant Rev. J. H----n. Marginal note on two lines: died at sea|informant daughter. | entry_id | vol2_p017_r004 |
name_raw | J---n C-----y | |
burial_date_raw | 14 March 18-- | |
age_raw | 6- | |
abode_raw | B------k | |
officiant_raw | Rev. J. H----n | |
marginalia_raw | died at sea|informant daughter | |
file_name | register-vol2-p017.jpg | |
uncertainty_flags | ["partial_name","illegible_date"] |
A dash stands in for each character that could not be read. J---n is a five-character name with a three-character gap, and 18-- is a year whose last two digits are lost. The field names, the source reference, and the flags all come from a real run; the record itself is invented.
Suitable collections
Burial registers and grave record sheets make up most of the work I take on. Estate papers, minute books, ledgers, and other handwritten or printed records are assessed the same way: I read a sample and tell you whether the collection suits the service.
Suitability comes down to the material. Some hands stay consistent across a whole volume and others change from page to page, and image quality matters as much as the writing itself: focus, resolution, lighting, and whether the pages lie flat. Ruled columns with repeated headers are faster to work with than free-form pages. A page of many short entries costs less per record than a page of long prose, and the number of fields you need affects the work as well.
Send ordinary pages along with the difficult ones. A sample drawn only from the clearest pages understates the work, and one drawn only from the worst overstates it, so the assessment needs to reflect the whole collection. If a collection falls outside what I do well, I will say so.
How it works
Assess samples
You send representative pages; I read them and come back with what the collection contains, how demanding it will be, and whether it is a good fit.
Agree scope and quote
Before any transcription begins, we settle the fields, the output format, the approach to review, anything to be excluded, and what delivery looks like. The quote follows from that scope.
Transcribe and review
Records are transcribed with TextHarvester-assisted processing and then reviewed according to the approach we agreed, with confidence scores and uncertainty flags alongside the text.
Deliver the dataset
You receive the agreed files. Any remaining uncertainty, and the limits of what the output supports, are set out in the handover.
Timing is agreed as part of the scope. A sample assessment does not commit either of us to the work.
Quality and limits
Every row carries a confidence score for each field, produced during transcription, and an uncertainty_flags entry naming what the process was unsure about. Records that fall below the confidence threshold are flagged for review, and in the delivered dataset those rows can be edited and marked as reviewed.
A confidence score records how unambiguous a reading was to the process. It says nothing about whether the record is historically correct, so treat the scores as a guide to where your own checking should go.
Where the text cannot be read, the gap is marked. One dash stands in for each estimated illegible character, so J---n is a name with a three-character gap and ----- is a word that could not be read at all. A pipe marks a line break inside a field. Nothing is invented to fill a gap, and bracketed guesses such as [?] are not used.
Completeness and consistency checks are a different thing from historical interpretation. The process can confirm that a date field is filled in, that its format is consistent, and that it agrees with the column heading above it. It cannot tell you that the date itself is wrong, or that two entries naming the same person refer to two people. That reading is work for your own domain experts, and I recommend it.
The review a collection receives is agreed at the start of the project, since it depends on the material, the fields, and what the dataset has to support. Accuracy is measured against a set of records transcribed by hand, and reported for your collection.
About Dónal
I am Dónal O'Tiarnaigh, a GIS consultant working from the south of Ireland, and I hold a postgraduate qualification in Geographical Information Systems from University College Cork. Transcription sits alongside my mapping and spatial-analysis work. Both take a difficult source and produce data that can be checked and defended later.
You can read more about my background and how I work, and the services page covers the rest of what I take on.
Questions
How much does it cost?
The quote follows the sample assessment. It depends on the volume, the legibility of the material, the number of fields you need, and how much review is involved. I have no standard rate card and will not quote before I have seen the pages.
Can you handle difficult handwriting?
Usually, and the sample is what settles it. Send the pages that worry you along with the ordinary ones. Where a word or a character cannot be read it is marked with a dash, so the dataset shows you where the uncertainty lies.
How long does it take?
Timing is agreed with the scope. The size and condition of the collection, and the review approach, both affect it, so it is settled alongside the quote.
What files will I receive?
CSV as standard, which opens directly in Excel, and JSON where a system needs it. Native Excel workbooks (.xlsx) are not produced. Every row keeps its source file_name, with page and row references for burial registers.
Can the results be mapped or published online?
Yes, as a separate piece of work. A dataset delivered as CSV can feed a map, a searchable website, or another system, and I can quote for that on its own. You can see the kind of thing in the project archive.
How are the records handled?
Images are processed locally and sent only to the AI provider configured for that project. Transcribed records are stored locally. If you would rather settle the handling arrangements before sending anything, ask and we can do that first.
Get in touch
Email a few representative pages, and tell me:
- your organisation or project
- the type of record: a burial register, grave record sheets, or something else
- the approximate number of pages or records
- the output you need, and what you want to do with it
- any deadline you are working to
A first enquiry can be sent without attachments. If material is restricted, describe it first and we can agree what is appropriate to send.
You can also reach me through the contact page.