Collaboration
Open to research collaborations, student projects, and mentoring.
What I work on
I work on LLM safety at EPFL DLab with Prof. Robert West. Most of my current work targets the pre-training phase: how safety, values, and personas form while a model is trained, and how to shape them from the first token instead of patching them afterwards. This is the Model Raising paradigm; Synthetic Persona Pretraining is our first large-scale instance of it.
Topics I am especially keen to collaborate on:
- Safety in pre-training: data curation, synthetic data, and alignment from token zero.
- Personas inside LLMs: how they emerge during training and how they connect to jailbreaks, refusals, and general behaviour.
- Interpretability of safety-relevant directions: persona vectors, refusal directions, steering, and monitoring.
- Safety evaluation: jailbreak and red-teaming benchmarks, LLM judges, and robust measurement.
- Uncertainty estimation and trustworthiness of LLMs, my earlier line of work, which I still enjoy collaborating on.
Working with me
I am happy to co-author, share code and data, and mentor. In 2026 I am a research mentor in SPAR, and I supervise MSc projects at EPFL. If you are an EPFL student, a researcher with an idea that fits the topics above, or an aspiring safety researcher looking for a mentor, get in touch. A short note on what you want to do and why is more useful than a full CV.
How to reach me
| vvmoskvoretskii@gmail.com (fastest) | |
| Telegram | @VityaVitalich |
| X | @vitya_vitalich |
| GitHub | VityaVitalich |
| Google Scholar | profile |
Students I have supervised
- Dominic Bazina-Grolinger (2026, EPFL) — MSc project: “Jailbreaking Monitoring with Persona Vectors”
- Wiktoria Rozkosz (2026, EPFL) — MSc project: “Jailbreaking with User-Based Persona”
- Edgar Desnos (2026, EPFL) — MSc project: “Multi-Agent Cooperation Simulations”
- Vladimir Ganzhara (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
- Gleb Stenin (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
- Vlad Knyazhevski (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
- Luiza Nigogosova (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
- Ekaterina Neminova (2023, HSE) — BSc thesis, co-supervised with Irina Nikishina; led to the ACL 2024 paper “TaxoLLaMA”
- Alina Lobanova (2023, HSE) — BSc thesis, co-supervised with Irina Nikishina; led to the ACL 2024 paper “TaxoLLaMA”