Collaboration

Open to research collaborations, student projects, and mentoring.

What I work on

I work on LLM safety at EPFL DLab with Prof. Robert West. Most of my current work targets the pre-training phase: how safety, values, and personas form while a model is trained, and how to shape them from the first token instead of patching them afterwards. This is the Model Raising paradigm; Synthetic Persona Pretraining is our first large-scale instance of it.

Topics I am especially keen to collaborate on:

  • Safety in pre-training: data curation, synthetic data, and alignment from token zero.
  • Personas inside LLMs: how they emerge during training and how they connect to jailbreaks, refusals, and general behaviour.
  • Interpretability of safety-relevant directions: persona vectors, refusal directions, steering, and monitoring.
  • Safety evaluation: jailbreak and red-teaming benchmarks, LLM judges, and robust measurement.
  • Uncertainty estimation and trustworthiness of LLMs, my earlier line of work, which I still enjoy collaborating on.

Working with me

I am happy to co-author, share code and data, and mentor. In 2026 I am a research mentor in SPAR, and I supervise MSc projects at EPFL. If you are an EPFL student, a researcher with an idea that fits the topics above, or an aspiring safety researcher looking for a mentor, get in touch. A short note on what you want to do and why is more useful than a full CV.

How to reach me

Emailvvmoskvoretskii@gmail.com (fastest)
Telegram@VityaVitalich
X@vitya_vitalich
GitHubVityaVitalich
Google Scholarprofile

Students I have supervised

  • Dominic Bazina-Grolinger (2026, EPFL) — MSc project: “Jailbreaking Monitoring with Persona Vectors”
  • Wiktoria Rozkosz (2026, EPFL) — MSc project: “Jailbreaking with User-Based Persona”
  • Edgar Desnos (2026, EPFL) — MSc project: “Multi-Agent Cooperation Simulations”
  • Vladimir Ganzhara (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
  • Gleb Stenin (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
  • Vlad Knyazhevski (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
  • Luiza Nigogosova (2024, HSE) — MS project: “Aligning LLM Confidence with Truthfulness with Self-Play”
  • Ekaterina Neminova (2023, HSE) — BSc thesis, co-supervised with Irina Nikishina; led to the ACL 2024 paper “TaxoLLaMA”
  • Alina Lobanova (2023, HSE) — BSc thesis, co-supervised with Irina Nikishina; led to the ACL 2024 paper “TaxoLLaMA”