Chuyển giọng nói thành văn bản (Speech-to-Text)
POST /audio/transcriptions
Nhận một file âm thanh và trả về nội dung lời nói dưới dạng văn bản. Request gửi dạng multipart/form-data.
Tham khảo
Chi tiết tham khảo tại tài liệu LiteLLM API
Cấu trúc yêu cầu (Form Data)
modelstringRequiredID của mô hình speech-to-text: gemini-3.5-transcribe-preview.
filefileRequiredFile âm thanh cần chuyển thành văn bản, ví dụ .mp3.
Cấu trúc Phản hồi (Response Body)
textstringNội dung lời nói trong file âm thanh.
usageobjectSố token đã dùng (input_tokens, output_tokens, total_tokens).
- curl
- Python (openai)
curl https://api.thucchien.ai/audio/transcriptions \
-H "Authorization: Bearer <your_api_key>" \
-F model=gemini-3.5-transcribe-preview \
-F file=@speech.mp3
from openai import OpenAI
client = OpenAI(
api_key="<your_api_key>",
base_url="https://api.thucchien.ai"
)
with open("speech.mp3", "rb") as audio_file:
transcript = client.audio.transcriptions.create(
model="gemini-3.5-transcribe-preview",
file=audio_file,
)
print(transcript.text)
Ví dụ phản hồi
{
"text": "Xin chào các đội thi.",
"usage": {
"type": "tokens",
"input_tokens": 51,
"output_tokens": 6,
"total_tokens": 57
},
"task": "transcribe"
}