หน้าแรกHome โมเดล AIAI Models

เครื่องเล่น Gemini 3.8: การสังเคราะห์เสียงที่ล้ำหน้า Gemini 3.8: A Leap in Text-to-Speech Innovation

โมเดล AIAI Models ข่าวNews 26 กันยายน 2569 26 September 2026 อ่าน 2 นาที 2 min read Oneable Team
Gemini 3.8: A Leap in Text-to-Speech Innovation

ลองสัมผัสความก้าวหน้าของ Gemini 3.8 ในการสังเคราะห์เสียงด้วยโมเดลที่ออกแบบใหม่จาก Google

Discover Gemini 3.8's advancement in text-to-speech with Google's newly designed models.

บทนำ

Google ได้เปิดตัวโมเดลสังเคราะห์เสียง Gemini 3.8 ตัวใหม่ ซึ่งประกอบด้วย gemini-3.8-flash-tts และ gemini-3.8-flash-lite-tts ที่รองรับเสียงกว่า 2,000 เสียง พร้อมความสามารถในการสร้างเสียงที่กำหนดเองได้จากตัวอย่างเสียงเพียง 30 วินาที

Gemini 3.8 ได้รับการยกย่องว่าเป็นอีกขั้นหนึ่งที่สำคัญในการสังเคราะห์เสียง เพื่อตรวจสอบตลาดที่มีการเติบโตอย่างรวดเร็ว วิวัฒนาการของเทคโนโลยีนี้มีความสำคัญอย่างไร เราจะพาคุณไปสำรวจฟังก์ชันและประโยชน์ที่มันสามารถให้ได้

คุณสมบัติและการทำงาน

Gemini 3.8 ถูกออกแบบมาเพื่อให้ง่ายต่อการนิยามบทสนทนาระหว่างตัวละครหลายตัว ซึ่งแต่ละตัวมีเสียงและสไตล์เสียงที่แตกต่างกัน เอพีไอนี้สามารถประยุกต์ใช้ได้อย่างหลากหลาย ตั้งแต่ออดิโอสำหรับความบันเทิงไปจนถึงการฝึกงานด้านการสื่อสาร

ตัวอย่างการใช้งานที่น่าสนใจคือตัวอย่างบทสนทนาระหว่างนกกระทุงสองตัวเกี่ยวกับการย้ายไปยัง Pacifica Pier ซึ่งใช้ Claude 4.5 Opus ในการเขียนสคริปต์และสร้าง URL เพื่อใช้เครื่องมือ Gemini ส่งผลให้ใช้เวลาเพียง 20 วินาทีในการสร้างออดิโอความยาว 1 นาที 18 วินาทีจาก Gemini 3.8 Flash TTS

เบื้องหลัง

Gemini 3.8 เป็นผลิตผลของ Google บริษัทซอฟต์แวร์ยักษ์ใหญ่ที่มีชื่อเสียงด้านปัญญาประดิษฐ์และการเรียนรู้ของเครื่อง โมเดลนี้ใช้กุญแจของ API ของ Gemini ที่เปิดทิ้งไว้ ทำให้สามารถสร้างสรรค์เสียงได้หลากหลายรูปแบบ ซึ่งมาพร้อมกับปริมาณเสียงที่มากกว่า 2,000 รายการ

Google ใช้ภาษาโปรแกรมขั้นสูงและอัลกอริทึมที่ซับซ้อนในการพัฒนาเทคโนโลยีนี้ สำหรับนักพัฒนาที่ต้องการเพิ่มความสามารถในการปรับให้เหมาะสม เครื่องเล่นนี้ก็คือแหล่งข้อมูลที่สำคัญ

ประโยชน์ของเครื่องเล่น

ผู้พัฒนาสามารถสร้าง URL ที่บันทึกการตั้งค่าเสียงต่าง ๆ ไว้ได้ง่าย ๆ เพื่อใช้ซ้ำหรือแบ่งปันกับผู้อื่น การสร้างเสียงคุณภาพสูงสามารถทำได้อย่างรวดเร็วและมีประสิทธิภาพในค่าบริการที่ถูกมาก

การรองรับการใช้งานหลายภาษาและความสามารถในการปรับแต่งเสียงตามใจชอบทำให้ Gemini 3.8 กลายเป็นเครื่องมือที่ทรงพลังในอุตสาหกรรมการสังเคราะห์เสียง

สรุป

ในยุคของการพัฒนา AI ที่เติบโตแบบก้าวกระโดด Gemini 3.8 ได้แสดงให้เห็นถึงขีดความสามารถในการสังเคราะห์เสียงที่น่าทึ่งและหลากหลาย ถือเป็นนวัตกรรมที่ยกระดับการสื่อสารและการเปิดโอกาสในการสร้างสรรค์ใหม่ๆ

ที่มา: Simon Willison — https://simonwillison.net/2026/Sep/23/gemini-tts-playground/

Introduction

Google has unveiled the new Gemini 3.8 text-to-speech models, featuring gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, with support for over 2,000 voices and the ability to create a custom voice using just a 30-second audio sample.

Gemini 3.8 is hailed as a significant leap in text-to-speech technology, transforming a rapidly growing market. We will explore its functions and benefits that enhance the capabilities of modern-day speech synthesis.

Features and Functionality

Designed to easily define conversations between multiple characters, each with distinct voices and voice style instructions, this API offers versatility for applications ranging from entertainment audio to workplace communication training.

An intriguing example is a simulated conversation between two pelicans debating a move to the Pacifica Pier, scripted using Claude 4.5 Opus and rendered via Gemini's toolkit, generating 1 minute 18 seconds of audio in just 20 seconds using the Gemini 3.8 Flash TTS.

Background

Gemini 3.8 is a product of Google, a renowned giant in the software industry, particularly in AI and machine learning. By leveraging an open CORS policy of the Gemini API, the tool allows creating various voice styles, supported by a library of over 2,000 audio samples.

Google employs advanced programming languages and sophisticated algorithms in developing this technology. For developers looking to enhance their projects, this playground serves as an essential resource.

Benefits of the Playground

Developers can easily generate bookmarkable URLs, preserving voice settings for repeat or shared use. High-quality speech synthesis can be achieved swiftly and cost-effectively.

Its multilingual support and ability to customize voices as desired make Gemini 3.8 a powerful tool in the realm of speech synthesis.

Conclusion

In the era of booming AI advancements, Gemini 3.8 showcases impressive and diverse capabilities in speech synthesis, presenting an innovation that elevates communication and opens up new creative avenues.

Source: Simon Willison — https://simonwillison.net/2026/Sep/23/gemini-tts-playground/

ที่มา:Source: simonwillison.net/2026/Sep/23/gemini-tts-playground/

เกี่ยวกับผู้เผยแพร่About the publisher

ผู้เขียนAuthor
Oneable Team
บริษัทCompany
Oneable — AI-Powered Software Development Agency
ความเชี่ยวชาญExpertise
LLM & RAG, AI Agent, Web/Mobile, MLOps
ติดต่อContact
www.oneable.co.th/contact

บทความที่เกี่ยวข้องRelated articles

DevFest 2026 กลับมาอีกครั้ง สนุกและเรียนรู้จากเหล่านักพัฒนาทั่วโลก DevFest 2026 Returns: Connect with Developers Worldwide

โมเดล AIAI Models ข่าวNews 23 ก.ย. 256923 Sept 2026 2 นาทีmin

รองรับรูปแบบเสียงที่ซับซ้อนด้วย Gemini 3.5 Transcribe Enhance Complex Audio with Gemini 3.5 Transcribe

โมเดล AIAI Models ข่าวNews 19 ก.ย. 256919 Sept 2026 2 นาทีmin

เครื่องมือปรับปรุงข้อความคอมมิต commit-rewriter 0.1 Introducing commit-rewriter 0.1 for Cleaner Git Commits

โมเดล AIAI Models ข่าวNews 18 ก.ย. 256918 Sept 2026 2 นาทีmin

AI กับการค้นพบโมเลกุลต้านจุลชีพใหม่ AI Innovations in Discovering New Antimicrobial Molecules

โมเดล AIAI Models อธิบายExplainer 17 ก.ย. 256917 Sept 2026 2 นาทีmin

วิธีการประเมินโมเดล LLMs ก่อนการใช้งานจริง Evaluating LLMs: Quick and Effective Methods

โมเดล AIAI Models How-toHow-to 16 ก.ย. 256916 Sept 2026

ฝึกโมเดลโค้ดให้สร้างภาพสีน้ำด้วย TRL และ OpenEnv Training Models to Create Watercolor Art with TRL & OpenEnv

โมเดล AIAI Models How-toHow-to 16 ก.ย. 256916 Sept 2026