About
Jan Scholtes
Prof. dr. ir. Johannes (Jan) C. Scholtes has spent more than forty years at the intersection of artificial intelligence and human language — long before either field became fashionable. His first publications date to 1987, when he built a knowledge-based optical character recognition system at TU Delft. His 1993 PhD thesis at the University of Amsterdam — Neural Networks in Natural Language Processing and Information Retrieval — prefigured by decades the architectures now at the centre of modern AI.
Since 2009 he holds the Endowed Chair of Text Mining at the Department of Advanced Computing Sciences at Maastricht University, where he teaches the Advanced NLP and Agentic Conversational Search courses and supervises graduate research spanning legal AI, clinical NLP, fraud detection, and agentic systems. He is a Fellow of the SIKS School of the Royal Netherlands Academy of Arts and Sciences (KNAW).
His applied work spans some of the most demanding professional environments in which language AI is deployed. Earlier in his career he served as an officer in the Dutch Naval Intelligence Service, an experience that shaped a lasting interest in the operational uses of AI in defence, security, and law enforcement contexts. As a long-standing authority in eDiscovery and LegalTech, he has worked with and for major law firms, auditing institutions, and multinational corporations on large-scale document review, fraud investigation, and regulatory compliance. His work on text mining for auditing was published in the European Court of Auditors Journal; he presented at the International Conference on Artificial Intelligence and Law (ICAIL) and at Leiden University’s Faculty of Law. He co-authored two books on the subject: Big Data Analytics for eDiscovery (Edgar Publishing, 2021) and The LegalTech Bridge (Meesterlijk, 2021).
On the industry side, Jan is Venture Partner at Endeit Capital, a leading European growth-equity fund focused on technology, and serves as Board Member at IPRally, the AI-native patent intelligence company whose search technology brings deep NLP to patent analytics at global scale. His chair at Maastricht University was established with support from ZyLAB Technologies, one of the world’s leading platforms for AI-powered review of large document collections, used by Fortune 500 companies, enforcement agencies, and courts worldwide.
His publication record runs to more than fifty peer-reviewed papers and book chapters — from the 1991 AAAI Spring Symposium and IEEE IJCNN to recent work on deep learning for clinical NLP, anomaly detection in complex networks, and conversational AI for assessment. He holds patents in text mining and information retrieval. A full list of publications is available at textmining.nu.
The posts on this site draw on that accumulated record: forty years of building systems that turn language into knowledge, and the conviction that the most important problems in AI today are not technical but organisational — who governs it, who controls the data, and whether the architecture is built to last.
The Name
The Latin name for Maastricht is Trajectum ad Mosam — “the crossing at the Meuse.” From Traject, add -orium as in auditorium or laboratorium, and you get Trajectorium: an institute for trajectories.
The double meaning is intentional. A trajectory is a path through space and time — the route an object follows, the course a system traces, the chain of decisions an AI agent works through from observation to action. The word bridges classical Latin and contemporary AI in a single name, grounded in the city where the work happens.
The metaphors run deeper than the etymology. The posts on this site argue that the most valuable trajectory in modern AI is the one from unstructured text — written by one generation of practitioners, for the benefit of the next — to training data for agents that carry that knowledge forward. A farming logbook written to train the next farmhand. A legal SOP written to train the next caseworker. A clinical note written for the next doctor on the ward. The knowledge was always meant to travel forward in time. What changed is who receives it at the other end.
That trajectory — from human-readable to machine-learnable, from institutional memory to sovereign, energy-efficient AI that runs on infrastructure an organization actually controls — is the practical and intellectual concern this site exists to map.