Skip to content
speechinfraConsole

Alibaba · Qwen · Speech alignment

Qwen3 Forced Alignment

Aligns an existing transcript to speech and returns precise word-level timestamps.

qwenasr-align
$0.005 / audio minuteTry in Console

Current model-specific price from our catalog. Requests require your project API key and sufficient balance. Availability is checked at execution time.

Input and capabilities

Upload a short audio or video file using multipart/form-data: at most 30 seconds and 20 MiB. Use WAV for the examples below. Container compatibility depends on the model.

  • transcription
  • word timestamps
  • alignment

Automatic language detection: not supported.

Supported language codes: zh-CN, en, yue, fr, de, it, ja, ko, pt, ru, es.

API endpoint

POST https://api.speechinfra.com/v1/inference/qwenasr-align

Authenticate with your Speechinfra project key. Read the quickstart and billing guide.

cURL
curl --fail-with-body "https://api.speechinfra.com/v1/inference/qwenasr-align" \
  -H "Authorization: Bearer $SPEECH_API_KEY" \
  -F "file=@clip.wav" \
  --form-string "language=en" \
  --form-string "transcript=Verbatim transcript matching the clip"
Python
# pip install httpx
import os
import httpx

headers = {"Authorization": "Bearer " + os.environ["SPEECH_API_KEY"]}
with open("clip.wav", "rb") as audio:
    response = httpx.post(
        "https://api.speechinfra.com/v1/inference/qwenasr-align",
        headers=headers, files={"file": ("clip.wav", audio, "audio/wav")},
        data={
    "language": "en",
    "transcript": "Verbatim transcript matching the clip"
}, timeout=90,
    )
response.raise_for_status()
print(response.json())
Node.js / fetch
// Node.js 22+, no SDK required.
import { readFile } from 'node:fs/promises';
const form = new FormData();
form.set('file', new Blob([await readFile('clip.wav')]), 'clip.wav');
for (const [key, value] of Object.entries({
  "language": "en",
  "transcript": "Verbatim transcript matching the clip"
})) form.set(key, value);
const response = await fetch("https://api.speechinfra.com/v1/inference/qwenasr-align", {
  method: 'POST', body: form, signal: AbortSignal.timeout(90000),
  headers: { Authorization: 'Bearer ' + process.env.SPEECH_API_KEY },
});
if (!response.ok) throw new Error(await response.text());
console.log(await response.json());

Request parameters

Generated from the API schema for this deployment. Conditional parameters apply only when their controlling setting is enabled.

ParameterTypeDetails
file

Required

string

Short audio or video file. Direct API limit: 30 seconds, 20 MiB.

languagestring

enum: zh-CN, en, yue, fr, de, it, ja, ko, pt, ru, es · default: "en"

All accepted values
[
  "zh-CN",
  "en",
  "yue",
  "fr",
  "de",
  "it",
  "ja",
  "ko",
  "pt",
  "ru",
  "es"
]
transcript

Required

string

Verbatim transcript of the uploaded clip.

max 20000 chars

Complete request schema
{
  "type": "object",
  "additionalProperties": false,
  "required": [
    "file",
    "transcript"
  ],
  "properties": {
    "file": {
      "type": "string",
      "format": "binary",
      "description": "Short audio or video file. Direct API limit: 30 seconds, 20 MiB."
    },
    "language": {
      "type": "string",
      "default": "en",
      "enum": [
        "zh-CN",
        "en",
        "yue",
        "fr",
        "de",
        "it",
        "ja",
        "ko",
        "pt",
        "ru",
        "es"
      ]
    },
    "transcript": {
      "type": "string",
      "minLength": 1,
      "maxLength": 20000,
      "description": "Verbatim transcript of the uploaded clip."
    }
  }
}

Response schema

The schema below describes the response; optional fields depend on model settings.

View complete response schema
{
  "$defs": {
    "Word": {
      "properties": {
        "word": {
          "title": "Word",
          "type": "string"
        },
        "start": {
          "title": "Start",
          "type": "number"
        },
        "end": {
          "title": "End",
          "type": "number"
        },
        "confidence": {
          "anyOf": [
            {
              "type": "number"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "title": "Confidence"
        },
        "speaker": {
          "anyOf": [
            {
              "type": "string"
            },
            {
              "type": "null"
            }
          ],
          "default": null,
          "title": "Speaker"
        }
      },
      "required": [
        "word",
        "start",
        "end"
      ],
      "title": "Word",
      "type": "object"
    }
  },
  "properties": {
    "words": {
      "items": {
        "$ref": "#/$defs/Word"
      },
      "title": "Words",
      "type": "array"
    },
    "language": {
      "title": "Language",
      "type": "string"
    },
    "model": {
      "default": "qwen3-forced-aligner-0.6b",
      "title": "Model",
      "type": "string"
    }
  },
  "required": [
    "words",
    "language"
  ],
  "title": "AlignmentResult",
  "type": "object"
}