- Experience
- 1+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 20 hours ago
- Work mode
- In office
- Education
- Master's degree
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
We are looking for a skilled engineer to advance the development and optimization of large-scale Text To Speech (TTS) models. This role involves enhancing model architectures, pre-training in-context learning (ICL), fine-tuning supervised fine-tuning (SFT), and engaging in related speech synthesis activities. The candidate will also be expected to stay abreast of the latest advancements in voice communication technologies and contribute to Baidu's speech synthesis projects deployed in multiple international products.
Responsibilities
- Optimize structures and overall efficiency of Text To Speech models.
- Engage in the development phases including pre-training and fine-tuning of TTS systems.
- Continuously track and integrate new research and innovation in voice technologies.
- Participate actively in implementing speech synthesis solutions for Baidu's overseas product portfolio.
Requirements
- Master's degree or higher in computer science or closely related field.
- At least one year of industry experience working with speech synthesis systems.
- Familiarity with Linux environments and proficient programming in Python.
- Experience with deep learning frameworks, specifically PyTorch.
- Strong communication skills; passionate about technology, eager to learn, and proactive in approach.
- Desirable: Publications in leading conferences or journals (e.g., Interspeech, ICASSP) or recognition in speech technology competitions.
Minimum education
Master's Degree
Skills
Tools & software
How they work
Communication
Initiative
Positive Attitude