What is your background, and how did you end up at the HCSS Datalab?
Going into university, I was pulled in two directions at once: towards computer science on one side, and towards social sciences on the other. For a long time, those felt like separate lives. I ended up studying both a BSc in Business Administration and a BSc in Information Sciences, and then an MSc in Data Science, at the University of Amsterdam. During my master I was exposed to the kind of problems you can state in a single sentence but that resist solution in non-polynomial time.
That was where I learned that even calculating the folded structure of a simplified amino acid chain smaller than insulin, where the pieces are known and the physics is understood, can still defeat our best algorithms. Alongside my studies I taught at the university’s Institute of Informatics, lecturing on everything from statistics and programming through to digital societies and emerging technological threats, and supervising student research along the way.
It was work I enjoyed, but academia left one appetite unfed: the international-relations and social-science side of me never had anywhere to go. The technical world and the human world stayed in separate rooms. Then HCSS crossed my path, and for the first time the two directions met in the same place. Now I’m here.

What exactly is the HCSS Datalab, and why was it created?
The Datalab is where we do quantitative, fact-based analysis to inform strategic decision-making. And strategic decision-making is, by its nature, exposed to risk. Uncertainty is what makes long-term choices so hard: it discourages bold action and keeps you from realising the impact you’re capable of. So, what we do in the Datalab is model complex phenomena, and try to track, quantify and compare the consequences of decisions, so that risk stops being a vague fog and becomes something you can reason about. Take horizon-scanning as an example: rather than waiting for a crisis to announce itself, we try to read the early signals across the defence, security and geopolitics domains and put numbers to them, so a decision-maker can see where the ground is shifting before it moves.
With data and deep subject-matter knowledge, we help sharpen the decision or recommendation itself and then tell the story in a way that reaches the people who need to hear it. What makes it work from my perspective is the mix of people: philosophers and historians who understand how people behave, international-relations scholars who understand how conflict takes shape, and mathematicians, econometricians and computer scientists who are good at building systems and solving technical puzzles.
What has changed in the Datalab since you became Manager?
A lot has changed, across methods, data and products alike. The obvious catalyst on the methods side is the explosion of AI over the past few years. But I’d push back gently on the idea that it caught us off guard, because we’ve been on that terrain since the early days of NLP at HCSS. I believe that head start mattered. It meant we could keep pace with each new development rather than scramble to catch up, and genuinely fold generative AI into our workflows rather than bolt it on at the end.
Beyond methods, we’ve invested heavily in building our own quantitative knowledge infrastructure. To give two examples, we now hold one of the largest collections of Dutch population-survey data in the Netherlands, and we scrape tens of thousands of documents, images and videos every day on a wide variety of topics, including Russia and critical raw materials. We’ve seen an explosion in the sheer number of data types and datasets we work with at HCSS, which is exactly what I always envisioned, because a broader information landscape gives you more angles to reason from.
And, those datasets no longer sit in isolation. They now feed our platforms and research products directly, most of them open-source and open-access, so the work can be re-used and compounds instead of being rebuilt each time. Over the last two years we’ve built a whole suite of these instruments, spanning the diplomatic, information, military and economic dimensions of geopolitical space, some open-source for the public, some closed-source for a range of Dutch ministries, the police, NATO and others.

How important is data for policy research today?
More important than it has ever been, and increasingly indispensable. My view is that the essence of all research should have an empirical data component, and policy research should honour that. At the same time, I’m not saying you should only base everything on data and models; I’m saying data belongs somewhere inside the decision-making process, not off to the side of it. There’s the classic line, “all models are wrong, but some are useful” which George Box used in a 1976 paper to make the point that no model is ever fully accurate, though a simpler one can still be genuinely useful if you apply it with judgement. But that was fifty years ago. Fifty years ago, we simply didn’t have the methodology we have today to capture the dynamics of the work we do. So, my answer would be that current technological progress finally lets us unlock the potential of data for policy research in a way that wasn’t possible for most of the field’s history. The models are still wrong, in Box’s sense, but they’ve never been this useful, and we’ve never been better placed to know when to trust them.
What is the added value of a Datalab?
It follows directly from that last point. We can now combine qualitative research methods with more quantitative, empirically grounded ones, and let each make the other better. A China or Russia expert’s read on why a state behaves the way it does becomes far more powerful when you can test it against event data at scale; and the data, in turn, is meaningless without someone who understands the context well enough to know what it’s actually saying.
From my perspective the value of a Datalab lies precisely in that meeting point: it’s where a domain expert’s intuition stops being an untested hunch and becomes something you can probe, quantify and either confirm or overturn, and where a dataset stops being a wall of numbers and starts telling you something about the world. At the same time, neither tradition has to defer to the other; they sharpen each other, and the analysis that comes out is stronger than either could produce alone. I’ll admit it’s still fairly unusual to have a Datalab in the world we operate in, which is exactly why I consider it such a privilege that HCSS has one, and part of what makes my work rewarding.

How do data science and AI complement traditional think-tank research?
That depends on what you count as ‘traditional’. Most people picture the written report, so let’s take that as the starting point. What’s great about data science and AI, is that you can approach the very same topics from a completely complementary angle, seeing patterns across thousands of documents that no single analyst could hold in their head at once. But you can also deliver the findings in completely different form factors. The same information we once presented only as a report, take our Strategic Monitor Police publications or the Progress Strategic Monitor as examples, can now also live as an interactive dashboard, an explainer video, a generated podcast, or a slide deck. The underlying process of research stays the same, the problem, the puzzle, the paper, the peer-reviewed publication, but the way we can package and present what we find has multiplied.
HCSS works on defence, security and geopolitics. What’s uniquely hard about data in these fields?
The classic problem is that the data you most want to work with often simply isn’t available. Nobody hands you a clean spreadsheet of another state’s intentions. We work mostly with open-source data and OSINT, building on those public flows and our own proprietary datasets rather than relying on anything privileged or classified. The deeper difficulty here is that in geopolitics you frequently don’t even know who all the actors are, let alone what drives them. It’s a strange inversion of the protein-folding work I came from: there, at least, you know your amino acids and their polarity. In geopolitics, you often don’t even know all the pieces and the actors themselves have agency, they bluff, they form alliances, they change their minds, they flip the board as you’re trying to read it. In a lot of fields, the pieces are at least laid out on the table; here you’re often still working out which pieces belong to the problem at all, which makes the measurement challenge genuinely different from most research domains.

People hear “data science” and think AI and machine learning. How much of the Datalab’s work is actually AI?
AI has become the reflexive answer to almost every question lately, so it’s worth separating the terms you mention, and since everyone draws the line a little differently, here’s mine. Data science and AI overlap, but neither sits fully inside the other. Data science, for me, is about extracting insight from data, statistics, cleaning, visualisation, domain knowledge. AI is about building systems that behave intelligently. And machine learning is the tool that sits in the overlap: a subfield of AI that data scientists reach for all the time.
At the Datalab we used to do mostly data science, and we’ve since moved largely toward AI. I’d split that AI work into two forms. The first is applying AI inside our products and platforms. Think of generative pipelines that scrape thousands of news outlets every single day to surface the red flags that could precede a geopolitical shock, or models that characterise how states behave and track how that behaviour shifts over time. The second is using AI to build the platforms themselves, mostly working with agentic AI and agents. I still remember my programming course at university, where a single stubborn bug could cost you days on Stack Overflow. That has changed completely. We’ve moved from writing code to designing code, and from there to designing whole systems.
At the same time, you have to keep reflecting on how you use it. I’m glad I come from a pre generative AI era, because writing myself, and losing days to my own ‘handwritten’ code hunting for a single error, taught me a great deal about conceptualisation and how things actually work under the hood. So, you stay aware of how you use AI: how it shapes your work, and how it shapes you, both as an individual and as an organisation. That’s exactly why, back in 2024, we already wrote our manual for the responsible use of generative AI. We’ve kept updating it ever since, and the most recent version, our code of conduct, came out recently.
The Datalab has a strong record of interns who become colleagues. How important are they?
Let me start by saying that from my experience we’ve always had great interns, and I’m consistently amazed by their dedication, their motivation, and the fresh ideas they bring through the Datalab’s door. They’re here for a learning experience, but the learning genuinely runs both ways, and we often learn as much from them.
My view on interning is that you’re a full part of the team from day one, because we want you to have the complete working experience, and that means we give assistant Data Scientists real ownership over their work. That can feel a little daunting at first, being handed something that matters, but it’s what makes them an indispensable part of the team, often as full-stack developers working across different research products and platforms. And.. It’s probably no coincidence that every current member of the Datalab was an intern here first.

What AI developments should people be watching?
On the large scale, we’ve seen the boom in large language models and the rise of OpenAI and, more recently, Anthropic’s Claude. But what I’m also following closely isn’t only the models in the abstract; it’s the moment they cross over into the real-world applications and become a strategic enabler. AI’s move into defence is one of the clearest examples, and the pace in 2024 and 2025 is hard to overstate: economists estimate that quality-adjusted AI production in the US grew at over 2,000 percent a year across those two years. On the capability side, METR’s work on task time-horizons suggests frontier models are roughly doubling in what they can do every hundred days or so. Put those two together, an economy compounding at that rate and a capability curve that steep, and you understand why defence establishments have stopped treating this as a research curiosity.
On the methods side, the thing I’d point people toward is the AI that goes into the technology itself, which is far less famous than the ‘ChatGPT’ or ‘Gemini’ but arguably more consequential, precisely because of what it enables. Yann LeCun more recent work is worth reading here, particularly on why today’s language models aren’t the whole story and why other architectures matter.
Professionally, I’m spending a lot of my own time on two such approaches. One is Generative Agent-Based Modelling, where instead of writing a single equation for how a system behaves, you simulate a whole population of interacting artificial agents, each following its own simple rules, and watch what collective dynamics emerge, much like the way a crowd or an alliance produces behaviour no single member ever intended. The other is Graph Neural Networks, which learn over the relationships between entities rather than treating each in isolation. That last point is why they’re so promising for us: geopolitics isn’t a list of countries, it’s a living web of connections between them.
What excites you most about the future? What’s your perspective on the Datalab in the next 1-2 years?
1-2 years is a long time, especially now, and if you look at how different the world, and the world of AI, is from just 1-2 years ago, you hesitate to make statements about the future, let alone predict anything. But generally, I think we’ve seen only the tip of the iceberg of what these technologies make possible. Here’s a ‘historical’ example to put in perspective: 1-2 years ago, some of the most advanced methods at the intersection of policy research and NLP were topic modelling, sentiment analysis and named entity recognition, which, in Box’s terms, were wrong and somewhat useful.
Today, at the HCSS Datalab, I’m confident in saying that we’re strong at mapping the “lay of the land” across defence, security and geopolitics in a quantitative way, and at horizon-scanning for the themes just beginning to surface. We’re far beyond those traditional methods now, compared to where we stood 1-2 years ago, but looking at how quickly they went from cutting-edge to table stakes, it makes you wonder what today’s frontier tools will look like in another 1-2 years, and how modest they’ll seem by then.
But mapping the terrain isn’t the same as understanding what generates it, and that’s the direction we’re most excited about at the Datalab. There’s a long-standing argument within our lab that international relations need its own “omics” revolution, its own AlphaFold moment, and I think that’s exactly right. Just as AlphaFold learned to infer a protein’s hidden structure from data nobody could read directly, we want to infer the hidden structures behind geopolitical events, the strategic postures, escalation thresholds and deterrence equilibria we can’t measure head-on but might learn to model. This is our (WIP) geomics project: treating geopolitics not as a pile of isolated incidents to catalogue, but as a system with an underlying generative grammar. It won’t be easy, states, unlike proteins, bluff and change their minds, but even partial progress on a problem this important is worth the attempt.
In the nearer term, over the next year, our analysis still focuses mostly on verbal and textual si, and even there, there’s still so much left to gain. Of all the unstructured data in the world, we capture only a sliver of it. What if we could capture not only the verbal and textual exchange, but the paralinguistic behaviour of state leaders at scale too, the pauses, the tone, the non-verbal signals that so often carry what someone actually means rather than what they say? That’s one of the frontiers I most want to explore next. And 1-2 years from now, my hope is that the Datalab has grown into the place where this kind of work is simply expected, where pairing deep domain knowledge with genuinely novel methods across defence, security and geopolitics is what we’re known for, and where the team is even better equipped to do it than we are today.




