Position Description
At Sonar, we are actively seeking an innovative Machine Learning Scientist to join our Data & AI team and pioneer the next generation of our code analysis engine. You will be at the forefront of applying cutting-edge AI and Large Language Model (LLM) techniques to the complex domain of source code. Your work will directly shape our products, pushing the boundaries of static analysis to help millions of developers write better, more secure code. If you are driven to solve real-world problems by turning state-of-the-art research into practical, high-impact solutions, this is the role for you.
What You Will Do
Spearhead Research & Innovation
- Stay at the forefront of ML, Deep Learning, and LLMs, focusing specifically on their application to the Software Development Lifecycle (SDLC).
- Identify novel opportunities to enhance our products by designing, prototyping, and validating novel ML models for source code analysis.
- Develop advanced AI models to resolve complex bugs, vulnerabilities, and code smells, surpassing traditional static analysis capabilities.
- Build LLM-powered features, including Retrieval-Augmented Generation (RAG) for contextual code analysis, fine-tuning models on proprietary codebases, and exploring agentic systems for automated code remediation.
- Engineer data pipelines to gather, process, and version massive code-centric datasets required for training and evaluating specialized models at scale.
Develop Advanced AI Models
- Design, prototype, and validate novel ML models that identify and resolve complex bugs, vulnerabilities, and code smells.
- Collaborate with engineering and product teams to integrate successful ML prototypes into Sonar's cutting-edge products, ensuring they meet the needs of our global user base.
Build LLM-Powered Features
- Develop and implement advanced LLM-based solutions, including fine-tuning strategies (e.g., LoRA, QLoRA), advanced prompt engineering, and working with vector databases and semantic search.
- Collaborate closely with engineering and product teams to integrate successful ML prototypes into Sonar's cutting-edge products, ensuring they meet the needs of our global user base.
Communicate and Evangelize
- Clearly articulate and document complex technical concepts and research findings to both technical and non-technical stakeholders.
Experience and Qualifications
Requirements
- Advanced academic background (Master’s or PhD) in Computer Science, Machine Learning, or a related quantitative field.
- Strong industry experience in machine learning, with a solid understanding of modern software engineering practices and tools.
- Strong programming skills in Python and hands-on experience with core ML/DL frameworks (e.g., PyTorch, TensorFlow, Hugging Face). Familiarity with Java is a plus.
- Proven experience in applied Machine Learning, with a strong focus on Natural Language Processing (NLP) or, ideally, Programming Language Processing (PLP).
- Hands-on experience with modern LLM architectures and techniques, such as fine-tuning strategies (e.g., LoRA, QLoRA), advanced prompt engineering, building and optimizing Retrieval-Augmented Generation (RAG) pipelines, and working with vector databases and semantic search.
- Experience with large-scale data processing frameworks and cloud infrastructure (e.g. AWS).
- Experience driving research projects from initial ideation to a demonstrable prototype with a high degree of autonomy.
- Excellent communication skills in English, with a talent for explaining complex scientific topics clearly and concisely.