Skip to main content

Chuyển giọng nói thành văn bản (Speech-to-Text)

Mô hình Speech-to-Text (STT) nhận một file âm thanh và trả về nội dung lời nói dưới dạng văn bản, hỗ trợ tiếng Việt.

Model được hỗ trợ (Gemini Enterprise Agent Platform (trước đây là Vertex AI)):

  • gemini-3.5-transcribe-preview

Model OpenAI: gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-transcribe, whisper-1. Riêng gpt-transcribe phải gửi response_format=json, xem Model OpenAI và DeepSeek.

Endpoint: POST /audio/transcriptions

Request gửi dạng multipart/form-data (upload file), không phải JSON.

curl https://api.thucchien.ai/audio/transcriptions \
-H "Authorization: Bearer <your_api_key>" \
-F model=gemini-3.5-transcribe-preview \
-F file=@speech.mp3

Kết quả trả về là JSON, nội dung lời nói nằm trong trường text:

{
"text": "Xin chào các đội thi.",
"usage": {
"type": "tokens",
"input_tokens": 51,
"output_tokens": 6,
"total_tokens": 57
},
"task": "transcribe"
}
Kết hợp với Text-to-Speech

Bạn có thể dùng file âm thanh sinh ra từ Text-to-Speech để thử nhanh endpoint này.