Groundwork Data Management

Foundations for long-term information

Hello, my name is Matthijs Breebaart. I build tools for organisations that deal with large, long-lived collections of structured information, and I help people make better use of them.

I spent the 2000s and 2010s deep in XML and the standards around it. My main focus was the document side of XML: mixed content, schema validation, metadata, transformations. I also got caught up in the semantic web, things like JSON-LD and Turtle. Those were interesting times.

There is more of my history on LinkedIn, though LinkedIn may ask you to sign up first.

In the 2020s XML went quiet and I worked on other things. But XML has not gone away. What has gone away, are some of the fundamental limitations that we had to deal with in the 2010s. Today, powerful LLMs can do things that would have taken an excessive, and therefore unobtainable, amount of manual work. The challenge is how to use this new capacity, because it can be a double-edged sword.

Answering this challenge is what led to Remastery. Remastery is a workbench for improving and converting large XML collections with extensive AI support. Nothing is overwritten and every change can be undone, so you can let the models do much more than you would otherwise dare. Work that would once have taken a team months can now be compressed into weeks. I invite you to have a look.

Remastery blends strong document foundations with specific AI interventions. The latter cannot do without the former. That is where the name groundwork data management comes from. It is about getting the basics right before diving in. It is also about improving your existing documents first, because that is your best guarantee that LLMs can extract more value out of them.

Remastery is one side of this puzzle. Sidestnd is the other side. It is a tool for working with metadata resources. Sidestnd bridges the traditional gap between resource management and resource usage by production systems. It is aimed at metadata resources with a long life. It was triggered by a simple question: how can we figure out in 2030 which version of a controlled vocabulary was used in 2020?

If you think your data management foundation could be better, or if you hold a collection that needs work, send me an email.