Multimodal ML Engineer
We're looking for a Multimodal ML Engineer to join White Circle, an AI Safety company building the policy enforcement and optimization layer for AI systems. Backed by $11M from senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, and DeepMind, White Circle processes 100M+ API calls monthly and runs its own LLMs in production.
You will
-
Train and fine-tune large-scale multimodal models (vision-language, audio, speech, video) from scratch and from pretrained checkpoints.
-
Design experiments, build multimodal data pipelines, and train MoE architectures.
-
Build alignment pipelines (SFT, DPO, GRPO), optimize for production (quantization, distillation, streaming), and deploy end-to-end.
-
Define evaluation metrics that actually matter for the product.
Requirements
-
3+ years training large-scale multimodal models.
-
Strong PyTorch and distributed training experience (DeepSpeed, FSDP).
-
Deep familiarity with multimodal architectures – LLaVA, Qwen-VL, InternVL, Audio Flamingo, Whisper, HuBERT, Conformer or similar.
-
Hands-on RLHF/alignment across modalities (GRPO, DPO, reward modeling).
-
Both audio and video experience required – sequence modeling for each, plus large-scale dataset curation and production inference optimization.
-
Relocation to Paris or London (hybrid) required.
Bonus
-
Audio signal processing fundamentals – spectrograms, mel features, noise reduction.
-
MoE architecture experience.
We offer
-
$100k–$250k/year salary + equity; higher figures can be negotiated.
-
Official employment, visa and relocation help.
Compensation: $100K – $250K • Higher figures and equity are negotiable
- • $100K – $250K • Higher figures and equity are negotiable
Find more English Speaking Jobs in France on Arbeitnow
This description is a summary provided by the original source, not the full posting. Sign up free to read the full listing →