Join Synthesia, the global leader in AI video platform technology, as a Staff Research Engineer. Our cutting-edge research department is driving the future of interactive avatars and video diffusion models. With a global presence in Europe and the US, Synthesia operates with over 90% of the Fortune 100 and over 60,000 businesses worldwide. Our latest funding round, valued at billion, underscores our commitment to innovation and excellence.
As a Staff Research Engineer, you will contribute to the development of leading-edge research solutions, focusing on avatar-centric interactive video diffusion models. Our team works on real-world challenges, ensuring our research directly impacts the lives of businesses and consumers. You will lead the adaptation of diffusion models to incorporate diverse conditioning signals, enhance streaming video quality, and improve lip-sync accuracy. You will collaborate closely with data teams to ensure the highest quality of datasets and contribute to the development of robust evaluation frameworks.
We are seeking an expert in machine learning, specifically in diffusion models and video diffusion models, with experience in PyTorch and modern ML frameworks. Your responsibilities will include streaming infinitely long video sequences, understanding audio and motion cues, and improving lip-sync accuracy and motion realism. You will also stay updated with the latest research in world models, interactive human/agent modeling, and related fields.
Key responsibilities include:
- Adapting diffusion models to incorporate diverse conditioning signals
- Developing methods for streaming infinitely long video sequences at real-time rates
- Working on the perceptual layer of interactive agents, including audio and reaction generation
- Improving lip-sync accuracy, motion realism, and overall visual quality in video diffusion models
- Building robust evaluation frameworks and test suites
- Collaborating closely with the data team to define data needs and ensure high-quality datasets
- Staying up to date with the latest research in AI models, interactive modeling, and related areas
We are looking for individuals who:
- Are comfortable owning and executing on the responsibilities listed above
- Have a strong background in ML and computer vision with relevant industry experience
- Have hands-on experience with diffusion models, ideally avatar-centric or video-focused
- Have experience with PyTorch and familiarity with modern ML frameworks and tooling
- Have strong Python engineering skills and a commitment to clean, maintainable research code
- Have an outcome-driven, detail-oriented, and motivated attitude towards pushing research into real product impact
- Are clear communicators of hypotheses, experiments, and results
What will set you apart from others:
- Experience with audio-conditioned video diffusion models and deep knowledge of recent video DiT architectures
- Demonstrated ability to own the full model development pipeline end to end, from data preparation to model design, training, and evaluation
- A strong publication record in areas such as world models, interactive agents, or video diffusion models
We are excited to offer a competitive compensation package, including salary, stock options, and bonuses. We offer a hybrid work setting with an office in London, Amsterdam, Zurich, Munich, or remote in Europe. We value our employees' work and have a culture of continuous learning and innovation. Our company culture is highly collaborative, and regular planning and socials are held at our hubs.
At Synthesia, we pride ourselves on being a people-first company. We operate with the belief that people should always come first. We are committed to AI ethics, safety, and security, and we believe in building the future of AI responsibly and ethically.
Are you ready to be part of our cutting-edge research team and help shape the future of AI video agents? Apply now!