« If we only fine-tune models on translated data, we are going to introduce biases from whatever language these datasets were originally available in.»
That is the problem Liam Duignan is working on. A research engineer at CEA, Liam joined us at POSAÏS 2026 That is the problem Liam Duignan is working on. A research engineer at CEA, Liam joined us at OpenLLM France.
The idea is straightforward: if we want to build AI models for use in France, translating datasets from another language is not enough.
« For an educational chatbot that is to be used in France, we need to be using real human data that has been written by real French professors and used by real students within France. » That means building datasets from the language, educational context and practices they are actually meant to serve.
Together with Asma Graiess, Matteo Van Ypersele, Jérôme Deshayes and Olivier Ferret, Liam is working on exactly that: native French data for fine-tuning and evaluating large language models for mathematics.
Because developing AI in Europe is also about developing the right data for it☝🏼. And, as Liam puts it, that work can help Europe and the rest of the world move in the right direction!