เครื่องเล่น Gemini 3.8: การสังเคราะห์เสียงที่ล้ำหน้า Gemini 3.8: A Leap in Text-to-Speech Innovation

ลองสัมผัสความก้าวหน้าของ Gemini 3.8 ในการสังเคราะห์เสียงด้วยโมเดลที่ออกแบบใหม่จาก Google
Discover Gemini 3.8's advancement in text-to-speech with Google's newly designed models.
บทนำ
Google ได้เปิดตัวโมเดลสังเคราะห์เสียง Gemini 3.8 ตัวใหม่ ซึ่งประกอบด้วย gemini-3.8-flash-tts และ gemini-3.8-flash-lite-tts ที่รองรับเสียงกว่า 2,000 เสียง พร้อมความสามารถในการสร้างเสียงที่กำหนดเองได้จากตัวอย่างเสียงเพียง 30 วินาที
Gemini 3.8 ได้รับการยกย่องว่าเป็นอีกขั้นหนึ่งที่สำคัญในการสังเคราะห์เสียง เพื่อตรวจสอบตลาดที่มีการเติบโตอย่างรวดเร็ว วิวัฒนาการของเทคโนโลยีนี้มีความสำคัญอย่างไร เราจะพาคุณไปสำรวจฟังก์ชันและประโยชน์ที่มันสามารถให้ได้
คุณสมบัติและการทำงาน
Gemini 3.8 ถูกออกแบบมาเพื่อให้ง่ายต่อการนิยามบทสนทนาระหว่างตัวละครหลายตัว ซึ่งแต่ละตัวมีเสียงและสไตล์เสียงที่แตกต่างกัน เอพีไอนี้สามารถประยุกต์ใช้ได้อย่างหลากหลาย ตั้งแต่ออดิโอสำหรับความบันเทิงไปจนถึงการฝึกงานด้านการสื่อสาร
ตัวอย่างการใช้งานที่น่าสนใจคือตัวอย่างบทสนทนาระหว่างนกกระทุงสองตัวเกี่ยวกับการย้ายไปยัง Pacifica Pier ซึ่งใช้ Claude 4.5 Opus ในการเขียนสคริปต์และสร้าง URL เพื่อใช้เครื่องมือ Gemini ส่งผลให้ใช้เวลาเพียง 20 วินาทีในการสร้างออดิโอความยาว 1 นาที 18 วินาทีจาก Gemini 3.8 Flash TTS
เบื้องหลัง
Gemini 3.8 เป็นผลิตผลของ Google บริษัทซอฟต์แวร์ยักษ์ใหญ่ที่มีชื่อเสียงด้านปัญญาประดิษฐ์และการเรียนรู้ของเครื่อง โมเดลนี้ใช้กุญแจของ API ของ Gemini ที่เปิดทิ้งไว้ ทำให้สามารถสร้างสรรค์เสียงได้หลากหลายรูปแบบ ซึ่งมาพร้อมกับปริมาณเสียงที่มากกว่า 2,000 รายการ
Google ใช้ภาษาโปรแกรมขั้นสูงและอัลกอริทึมที่ซับซ้อนในการพัฒนาเทคโนโลยีนี้ สำหรับนักพัฒนาที่ต้องการเพิ่มความสามารถในการปรับให้เหมาะสม เครื่องเล่นนี้ก็คือแหล่งข้อมูลที่สำคัญ
ประโยชน์ของเครื่องเล่น
ผู้พัฒนาสามารถสร้าง URL ที่บันทึกการตั้งค่าเสียงต่าง ๆ ไว้ได้ง่าย ๆ เพื่อใช้ซ้ำหรือแบ่งปันกับผู้อื่น การสร้างเสียงคุณภาพสูงสามารถทำได้อย่างรวดเร็วและมีประสิทธิภาพในค่าบริการที่ถูกมาก
การรองรับการใช้งานหลายภาษาและความสามารถในการปรับแต่งเสียงตามใจชอบทำให้ Gemini 3.8 กลายเป็นเครื่องมือที่ทรงพลังในอุตสาหกรรมการสังเคราะห์เสียง
สรุป
ในยุคของการพัฒนา AI ที่เติบโตแบบก้าวกระโดด Gemini 3.8 ได้แสดงให้เห็นถึงขีดความสามารถในการสังเคราะห์เสียงที่น่าทึ่งและหลากหลาย ถือเป็นนวัตกรรมที่ยกระดับการสื่อสารและการเปิดโอกาสในการสร้างสรรค์ใหม่ๆ
ที่มา: Simon Willison — https://simonwillison.net/2026/Sep/23/gemini-tts-playground/
Introduction
Google has unveiled the new Gemini 3.8 text-to-speech models, featuring gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, with support for over 2,000 voices and the ability to create a custom voice using just a 30-second audio sample.
Gemini 3.8 is hailed as a significant leap in text-to-speech technology, transforming a rapidly growing market. We will explore its functions and benefits that enhance the capabilities of modern-day speech synthesis.
Features and Functionality
Designed to easily define conversations between multiple characters, each with distinct voices and voice style instructions, this API offers versatility for applications ranging from entertainment audio to workplace communication training.
An intriguing example is a simulated conversation between two pelicans debating a move to the Pacifica Pier, scripted using Claude 4.5 Opus and rendered via Gemini's toolkit, generating 1 minute 18 seconds of audio in just 20 seconds using the Gemini 3.8 Flash TTS.
Background
Gemini 3.8 is a product of Google, a renowned giant in the software industry, particularly in AI and machine learning. By leveraging an open CORS policy of the Gemini API, the tool allows creating various voice styles, supported by a library of over 2,000 audio samples.
Google employs advanced programming languages and sophisticated algorithms in developing this technology. For developers looking to enhance their projects, this playground serves as an essential resource.
Benefits of the Playground
Developers can easily generate bookmarkable URLs, preserving voice settings for repeat or shared use. High-quality speech synthesis can be achieved swiftly and cost-effectively.
Its multilingual support and ability to customize voices as desired make Gemini 3.8 a powerful tool in the realm of speech synthesis.
Conclusion
In the era of booming AI advancements, Gemini 3.8 showcases impressive and diverse capabilities in speech synthesis, presenting an innovation that elevates communication and opens up new creative avenues.
Source: Simon Willison — https://simonwillison.net/2026/Sep/23/gemini-tts-playground/
ที่มา:Source: simonwillison.net/2026/Sep/23/gemini-tts-playground/
เกี่ยวกับผู้เผยแพร่About the publisher
- ผู้เขียนAuthor
- Oneable Team
- บริษัทCompany
- Oneable — AI-Powered Software Development Agency
- ความเชี่ยวชาญExpertise
- LLM & RAG, AI Agent, Web/Mobile, MLOps
- ติดต่อContact
- www.oneable.co.th/contact