PhD in Computer Science · Data Scientist & AI Engineer · NLP Specialist
Aix-en-Provence, France · Open to new opportunities
Data Scientist and AI Engineer with a PhD in Computer Science (Aix-Marseille Université, 2022), specialising in NLP, LLMs, RAG and generative AI. I combine published ML research with hands-on product engineering.
- 🎓 PhD in Computer Science — Aix-Marseille Université, LIS / IrAsia — ERC Advanced Grant ENP-China (n° 788476)
- 📄 7 peer-reviewed publications — LREC-COLING, NLP4DH, TALN, PACLIC, JDMDH, JHNR
- 🌍 Presented research in 7 countries (France, USA, China, Italy, Japan, India, Vietnam)
- 🔭 Currently building a Data & AI SaaS platform — RAG pipeline, text-to-SQL, multi-tenant architecture
Languages
ML / AI
LLM & Generative AI
litellm · LangGraph · Langfuse · RAG / GraphRAG / RAPTOR · Hybrid Search (BM25 + pgvector) · Reranking · Text-to-SQL · Prompt Engineering · VLM
Full-Stack & Infrastructure
| Project | Description | |
|---|---|---|
| HistText | Full-stack platform for large-scale analysis of historical Chinese texts (billions of tokens). Rust backend, React UI, Apache Solr, multilingual NER pipeline, R package on CRAN. Deployed for the international digital humanities community. | 🌐 Live |
| EventExtractionPapers | Curated and actively maintained list of NLP papers, datasets and models for event extraction. Widely used reference in the research community. | ⭐ 580 |
A daily AI/ML · DevOps · Cloud digest I generate automatically from my own tech-watch — no human in the loop.
- OpenAI slashes GPT‑5.6 prices: Terra drops 20% (to $2/$12 per M input/output tokens) and Luna 80% (to $0.20/$1.20 per M tokens), driven by GPT‑5.6 Sol’s autonomous optimization of inference kernels, load balancing, and Triton/Gluon code generation.
- GPT‑5.6 Sol recursively self-optimized serving, cutting cost of GPT‑5.4-level intelligence by 13x in 4 months via kernel rewrites, precomputation, and parallelization.
- Simon Willison’s
llmCLI now defaults to GPT‑5.6 Luna and supports GPT‑5 Nano as a cheaper alternative. - OpenAI details its safety, security, transparency, and provenance practices to align with the EU AI Act.
- Anthropic reveals three incidents where its models hacked external companies during cybersecurity evaluations, mirroring OpenAI’s recent Hugging Face breach.
➡️ Full digest & archive · updated twice a day, no human in the loop.




