Medieval alchemists spent their lives trying to turn lead into gold. They never managed it. But the way they worked has a lot in common with how data teams build platforms today.
I named this site after their furnace, the athanor, so I want to start by explaining the link.
A furnace built for slow, steady heat
An athanor was designed to keep a low, even heat going for weeks at a time. Alchemists believed the work could not be rushed. Too much heat ruined the batch, and too little did nothing.
A data pipeline has the same job. When it runs every night without trouble, nobody notices it. But every report, dashboard and AI model depends on it, and when it stops, everything after it stops too.
Three stages, then and now
Alchemists described their work in stages, with the material getting purer at each step. A lakehouse does something similar with bronze, silver and gold layers.
Bronze is the raw material. The alchemist started with crude ore. In a data platform, bronze is where raw data lands: source extracts, API responses and files, kept exactly as they arrived. Nothing gets cleaned here.
Silver is where you clean and combine. The alchemists had a saying, solve et coagula, which means “dissolve and bring together.” You break the material down, remove what doesn’t belong, and put it back together in a better form. In the silver layer, that means removing duplicates, fixing data types, standardising codes and matching customers and products across systems.
Gold is what the business uses. Only a small amount of gold came out of a lot of ore. Gold tables are the same: smaller, cleaner and built for a specific use, such as revenue by region, a customer 360 view, or the features a machine learning model trains on.
Five lessons from the alchemists
1. Keep the raw material.
If a batch went wrong, an alchemist needed to know what went in. Your bronze layer does that job. When someone asks why a number in a gold table changed, the answer is usually in the raw data. If you’ve overwritten it, you can’t find out.
2. Steady and small beats big and risky.
Small loads that run every day are easier to manage than one huge monthly rebuild. When a source system changes, a small load breaks in a small way that’s easy to fix.
3. Count what you lose at every step.
Alchemists weighed their material before and after each stage, so they knew exactly how much was lost. Data teams should do the same.
I learned this on a recent project that matched aircraft maintenance records to the spare parts they use. The results looked thin, and everyone assumed the part-matching logic was the problem. So instead of tuning the matching, I counted the records at every step of the pipeline, from the raw tables to the final output.
The matching turned out to be fine. Most records were being lost long before they got there. One filter quietly threw away thousands of records because of how a single field was coded. A chain of inner joins dropped any record that was missing even one reference. Nothing failed and no error was raised. The rows just disappeared.
Once we could see the losses stage by stage, the conversation changed. The question was no longer “how do we match better?” but “why do most of our records never reach the matching step?” That’s a much more useful question, and it only came from counting.
4. Test before you publish.
Gold was tested with a touchstone before anyone paid for it. Data should be tested too. Check row counts, missing values, freshness and totals against the source before a gold table goes live, not after someone spots a wrong number in a meeting.
5. Not every dataset needs to be gold.
Most ore never became gold, and that was expected. Some data is only used for one-off analysis and can stay in silver. Building gold tables that nobody uses is wasted effort.
Where AI fits in
AI agents and RAG systems are only as good as the data they read. If a language model reads messy raw data, it will give confident answers built on duplicates and old records. Give it clean, well-defined gold data and the answers get much better.
What’s next
On this site I write about things I’ve actually built: AI agents, MCP servers, and tool comparisons across Snowflake, Databricks, AWS, Azure and GCP. Each post includes what worked, what broke and what I’d change.
If you’d like to get the next post by email, subscribe below. If you’re working on a data or AI problem of your own, you can reach me through the Contact page.
Leave a Reply