Deskripsi Pekerjaan
Informasi lengkap tentang posisi dan persyaratan
Ringkasan Yukerja
Lowongan Research Scientist, Text-to-Speech di Oddin kami kurasi dari Himalayas (kategori Teknologi & IT). Posisi ini ditandai sebagai remote — pastikan timezone dan syarat lokasi kandidat di deskripsi resmi. Yukerja.com bukan pemberi kerja — lamaran diproses di situs sumber resmi.
What you will be doing
- Research and train fast and quality SOTA TTS models for realistic and emotional voice generation for entertainment and education applications.
- You will be experimenting with different architectures / data to improve the quality and speed of the TTS model(s) and put the best results to production.
- Staying up to date with current research and coming up with new ideas / what to improve is very important for us!
- You will be in immediate collaboration with a team of 3 researchers specializing in TTS, and the product is supported by engineering and hardware stuff to ensure deployment
Skills you need
- Experience with training some text-to-speech / voice cloning models
- Solid knowledge of transformers, diffusion models, GANs
- Understanding of human speech and audio processing (sampling, spectrograms, vocoders)
- Proficiency in Python and key libraries (e.g., PyTorch, Hugging Face Transformers).
- Ability to keep up to date with research, understand papers, implement approaches; strong ML fundamentals and critical thinking
- Familiarity with modern speech synthesis models (GPT-based, flow matching… such as Vevo, StyleTTS, IndexTTS, Maskgct etc.)
- Contributions to open-source AI tools or research publications in Speech processing field
- Familiarity with AWS / similar clusters
Nice-to-have:
Originally posted on Himalayas