UWebASR

UWebASR — Speech to Text

Department of Cybernetics, University of West Bohemia · LINDAT/CLARIAH‑CZ
LINDAT/CLARIAH-CZ DARIAH-EU CLARIN.EU

The UWebASR service offers an accessible way to convert speech into text directly in your browser. It is designed to handle recordings of varying quality and length, providing fast results without the need to install any software. The system is developed by the Department of Cybernetics and NTIS research centre at the University of West Bohemia, as part of the LINDAT/CLARIAH‑CZ research infrastructure. UWebASR supports both contemporary and historical speech recognition tasks, making it useful for everyday applications as well as for research in the humanities and social sciences.

We now also offer a dedicated UWebASR Skill for agentic AI: https://github.com/honzas83/uwebasr-skill.

This service is intended exclusively for scientific and non-commercial use. For other types of use and detailed information about the service, please contact the authors.

1 · Select the language and recognition model

Idle

Available languages: Czech, Slovak, English, German, Dutch, Polish, Hungarian, Croatian, Serbian. For each language, multiple recognition models are provided. The newer Zipformer models include generic speech models and adapted oral history models, while the older Wav2Vec 2.0 models are kept for scientific reproducibility and long-term compatibility.

2 · Upload audio file

Click to select a file or drag & drop it here.

Common formats supported (WAV, MP3, M4A, MP4, …).
Recognition starts automatically after file selection.

3 · Result & Download

 

About

The UWebASR service is developed and maintained by the Department of Cybernetics, Faculty of Applied Sciences, University of West Bohemia in Pilsen.

It is also integrated into the national research infrastructure LINDAT/CLARIAH-CZ, which is part of the European CLARIN ERIC network. LINDAT/CLARIN’s mission is to provide open-access to digital linguistic data, tools, and services for the broad research and education community, while ensuring long-term preservation and sustainable operation.

More about LINDAT/CLARIN: lindat.cz

Data Privacy

For all recognition models listed on this page (accessed via the webpage or the HTTP API with the lindat/ prefix), the service is configured with a strict privacy policy that retains only anonymous session statistics. Audio recordings and generated transcriptions are processed entirely on-the-fly in system memory (RAM) and immediately discarded once the session finishes; no audio files or transcripts are stored, cached, or permanently archived on the servers, nor are they used for training models or any other purpose.

The underlying SpeechCloud platform runs entirely on physical servers located in the datacentre of the Faculty of Applied Sciences (Technická 8, Plzeň, Czech Republic). Audio, transcriptions, and other recognition data are not shared outside this infrastructure.

This website uses Google Analytics to collect website traffic statistics, such as page visits. Analytics collection is disabled before a selected audio file is processed. Website analytics are separate from speech recognition processing; audio recordings and generated transcriptions are not sent to Google Analytics. For further information or inquiries, please contact the service representative, Jan Švec (honzas@fav.zcu.cz).

Model performance

Model 🇨🇿 Czech 🇩🇪 German 🇸🇰 Slovak 🇬🇧 English 🇵🇱 Polish 🇭🇺 Hungarian 🇭🇷 Croatian 🇷🇸 Serbian
Whisper-large-v3 19.78 25.94 22.02 18.83 22.81 30.92 11.96 15.82
UWebASR (Zipformer) 14.02 13.97 10.28 11.46 15.68 16.35 4.60 8.02
UWebASR (Adapted Zipformer) 7.32 12.40 10.28 10.17 15.57 16.64

Numbers represent Word Error Rate (WER) in % (lower is better). Evaluation was conducted on a rich mix of accessible datasets for each language across diverse domains, including some datasets that were probably not part of Whisper’s training data.

Citing UWebASR

For all Zipformer models and for the general description of the service, please cite:

ŠVEC, Jan; LEHEČKA, Jan a IRCING, Pavel. Current State of the UWebASR - Web-Based ASR Service for Czech, Slovak, German, and English. CLARIN, 2025. ISSN 2773-2177.

@article{svec2025uwebasr,
  author={Jan {\v S}vec and Jan Lehe{\v c}ka and Pavel Ircing},
  title={Current State of the {UWebASR} - Web-Based {ASR} Service for {C}zech, {S}lovak, {G}erman, and {E}nglish},
  journal={CLARIN},
  year={2025},
  issn={2773-2177}
}

The older English, German and Czech models (Wav2Vec 2.0) should be cited as follows:

Lehečka, J., Švec, J., Psutka, J.V., Ircing, P. (2023) Transformer-based Speech Recognition Models for Oral History Archives in English, German, and Czech. Proc. INTERSPEECH 2023, 201-205, doi: 10.21437/Interspeech.2023-872

@inproceedings{lehecka23_interspeech,
  author={Jan Lehečka and Jan Švec and Josef V. Psutka and Pavel Ircing},
  title={Transformer-based Speech Recognition Models for Oral History Archives in English, German, and Czech},
  booktitle={Proc. Interspeech 2023},
  pages={201--205},
  year={2023},
  doi={10.21437/Interspeech.2023-872}
}

The older Slovak model (Wav2Vec 2.0) should be cited as follows:

Lehečka, J., Psutka, J.V., Psutka, J. (2023) Transfer Learning of Transformer-Based Speech Recognition Models from Czech to Slovak. In: Text, Speech, and Dialogue. TSD 2023. Lecture Notes in Computer Science, vol 14286. Springer, Cham. https://doi.org/10.1007/978-3-031-40498-6_29

@inproceedings{lehecka23_tsd,
  author={Jan Lehečka and Josef V. Psutka and Josef Psutka},
  title={Transfer Learning of Transformer-Based Speech Recognition Models from Czech to Slovak},
  booktitle={Text, Speech, and Dialogue},
  year={2023},
  publisher={Springer},
  pages={328--338},
  isbn={978-3-031-40498-6}
}

UWebASR HTTP API

The UWebASR HTTP API transcribes uploaded audio and remote audio sources. Use an HTTP POST request to upload audio data, or an HTTP GET request to process an audio URL. GET requests can also recognize a continuous live stream. Select the response format with the format query parameter; available formats include plain text, JSON, XML, and WebVTT. Results are returned progressively as they are recognized, except for TRS formats, which require the complete result.

Endpoint and model selection

Send requests to an endpoint in the following form:

https://uwebasr.zcu.cz/api/v2/lindat/{app_id}

The app_id path parameter selects the speech recognition model. Replace it with one of the model IDs listed below.

Available model IDs

Zipformer architecture (2023+)
Adapted Zipformer (2026+)
Wav2Vec 2.0 architecture (2020+)
Wav2Vec 2.0 architecture (2020+, oral histories, deprecated)

For example, the endpoint for the generic Czech Zipformer model is:

https://uwebasr.zcu.cz/api/v2/lindat/generic/cs/zipformer

Availability and limits

Recognition runs on a pool of SpeechCloud workers. If all workers for the selected model are busy, the API responds with HTTP status 503.

A recognition session can run for at most 3600 seconds of wall-clock time. Because recognition usually runs faster than real time, an audio file longer than one hour may still complete within this limit.

Request methods

GET: remote files and live streams

Pass the source audio URL in the required url query parameter. UWebASR downloads and recognizes the source using the SpeechCloud HTTP API User-Agent. Add stream=1 for a live source that should be recognized until the source or client connection closes. Without stream=1, recognition stops after the first result. Source and recognition errors are reported in the streamed response.

POST: upload audio

Send the audio bytes in the request body directly to the selected model endpoint. UWebASR accepts audio formats supported by FFmpeg and decodes the data progressively while they are uploaded.

HTTP Headers

The HTTP API response has the following additional headers:

HTTP GET/POST with output format specification

The output format is specified by the parameter format in the request:

https://uwebasr.zcu.cz/api/v2/lindat/app_id?format=webvtt

Supported formats

Plaintext (format=plaintext)

Clear plaintext in UTF-8 encoding:

Content-Type: text/plain; charset=UTF-8
Transcriber XML (format=trs) [does not support streaming]

Transcriber-accepted XML-based format. The output TRS file contains the file name, it can be passed to the API using the header Content-Disposition:

Content-Type: text/xml

Content-Disposition: filename="foo.wav"
Extended Transcriber XML (format=extended_trs) [does not support streaming]

XML Transcriber extended by confidence, the rest is the same as the TRS format.

Content-Type: text/xml
WebVTT (format=webvtt)

Web captions in WebVTT format with timestamps and transcript support, without confidence.

Content-Type: text/webvtt
JSON (format=json)

A JSON array containing objects with timestamps, confidence, and word fields:

The individual JSON objects corresponding to the recognized words are written as one JSON object per line forming a valid JSON array.

Example output (pretty-printed, does not have one JSON object per line):

Content-Type: application/json

[
{"start": 0.05999999865889549,
"end": 0.6299999859184027,
"word": "good",
"confidence": 0.9723399117549626,
"speech_end": false
},
{"start": 0.6299999859184027,
"end": 1.1099999751895666,
"word": "day",
“confidence": 0.9999999999999977,
"speech_end": true
}
]
SpeechCloud JSON (speechcloud_json)

Internal format containing all messages from the SpeechCloud platform. Suitable for further processing and integration into other platforms (it may also contain NLU results, etc.).

Content-Type: application/json; charset=UTF-8

The JSON file consists of a list of messages. Each message has a type property indicating the type of the message. The correct sequence of message types is the following:

If there is an error during the processing, the message {"type": "sc_error"} will appear in the message stream.

If the input is not processed till its end, the asr_input_processed will not appear in the message stream.

Error states

The errors which are encountered are indicated in the response body according to a given output format. The errors occurring before writing the HTTP response are indicated also using the HTTP status codes and HTTP status line (together with the indication in the response body).

Error codes

The error codes could be divided into several groups:

Example

An HTTP GET requesting recognition of a URL returning 404 (http://google.com/foo):

https://uwebasr.zcu.cz/api/v2/lindat/app_id?format=plaintext&stream=1&url=http://google.com/foo

returns HTTP status code 404 Not Found as provided by the up-stream server handling the URL (google.com). The error code is also included in the output HTTP response according to the required output format:

plaintext

The HTTP status code and the reason are specified on the new line in the output after the # symbol:

# 404 Not Found
trs, extended_trs

The HTTP status code and the reason are stored in an ErrorCode and ErrorReason elements in the output:

<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE Trans SYSTEM "trans-14.dtd">
<Trans audio_filename="input.wav">
  <Episode>
    <ErrorCode>404</ErrorCode>
    <ErrorReason>Not Found</ErrorReason>
  </Episode>
</Trans>
webvtt

The errors in WebVTT format are reported as comments after the NOTE label at the beginning of the line:

WEBVTT
            NOTE Error 404 Not Found
                
json, speechcloud_json

The error code and status are reported in a JSON object embedded into the output JSON array of the HTTP response:

[
{"status_code": 404, "status_reason": "Not Found"}
]

Examples

The following examples use curl command-line utility as a common HTTP client available almost everywhere.

To recognize the audio file test_wav/i_i2.wav in local directory:

curl -X POST -N \
--data-binary @test_wav/i_i2.wav \
'https://uwebasr.zcu.cz/api/v2/lindat/malach.cs?format=plaintext'

To recognize live audio stream from (http://icecast8.play.cz/cro1-128.mp3) into a JSON format use (note the stream=1 parameter):

curl -N "https://uwebasr.zcu.cz/api/v2/lindat/generic/cs/zipformer?format=json&stream=1&url=http://icecast8.play.cz/cro1-128.mp3"

Python script uwebasr.py

For Python integrations, a helper script is available: uwebasr.py from the skill repository.

Shell script uwebasr.sh

You can use the following convenience shell script for automatic processing of input files. It depends on ffmpeg and curl utilities installed in your system. It processes the input into the SpeechCloud JSON format and then it converts it into TXT and VTT formats. The files with .s.txt and .s.vtt suffixes contain also the information about short and long pauses.

#!/bin/bash
set -o nounset
set -o errexit
set -o pipefail

LANG=${1:?Please, pass LANG as $1}

URL="https://uwebasr.zcu.cz/api/v2/lindat/malach/${LANG}"
CONV_URL="https://uwebasr.zcu.cz/utils/v2/convert-speechcloud-json"

shift
x=${1:?Please, specify one or more input files}

for INPUT_FILE in "$@"; do
    JSON_FILE=${INPUT_FILE%.*}.json
    TXT_FILE=${INPUT_FILE%.*}.txt
    STXT_FILE=${INPUT_FILE%.*}.s.txt
    VTT_FILE=${INPUT_FILE%.*}.vtt
    SVTT_FILE=${INPUT_FILE%.*}.s.vtt

    echo "=== Recognizing to raw JSON: $JSON_FILE"
    ffmpeg -hide_banner -loglevel error -i "$INPUT_FILE" -ar 16000 -ac 1 -q:a 1 -f mp3 - |\
        curl --http1.1 --data-binary @- "${URL}?format=speechcloud_json" > "$JSON_FILE"
    echo "=== Converting to plaintext: $TXT_FILE"
    curl --data-binary "@${JSON_FILE}" "${CONV_URL}?format=plaintext" > "$TXT_FILE"
    echo "=== Converting to sentext: $STXT_FILE"
    curl --data-binary "@${JSON_FILE}" "${CONV_URL}?format=plaintext&sp=0.3&pau=2.0" > "$STXT_FILE"
    echo "=== Converting to WebVTT: $VTT_FILE"
    curl --data-binary "@${JSON_FILE}" "${CONV_URL}?format=webvtt" > "$VTT_FILE"
    echo "=== Converting to SentVTT: $SVTT_FILE"
    curl --data-binary "@${JSON_FILE}" "${CONV_URL}?format=sentvtt&sp=0.3&pau=2.0" > "$SVTT_FILE"
done
                

Shell script usage

To recognize the file test_wav/i_i2.wav using the Czech model, simply use:

uwebasr.sh cs test_wav/i_i2.wav

The output should look like:


=== Recognizing to raw JSON: test_wav/i_i2.json
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
100 56365    0  1204  100 55161     56   2604  0:00:21  0:00:21 --:--:--   312
=== Converting to plaintext: test_wav/i_i2.txt
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
100  1302    0    98  100  1204    971  11938 --:--:-- --:--:-- --:--:-- 14152
=== Converting to sentext: test_wav/i_i2.s.txt
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
100  1307    0   103  100  1204   1029  12028 --:--:-- --:--:-- --:--:-- 14206
=== Converting to WebVTT: test_wav/i_i2.vtt
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
100  1402    0   198  100  1204   1879  11430 --:--:-- --:--:-- --:--:-- 14604
=== Converting to SentVTT: test_wav/i_i2.s.vtt
  % Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed
100  1413    0   209  100  1204   2024  11663 --:--:-- --:--:-- --:--:-- 15031
                

You can also pass more than one file:

uwebasr.sh cs test_wav/input1.wav test_wav/input2.wav test_wav/input3.wav

For other languages, change the first parameter (en, de, cs, sk):

uwebasr.sh en test_wav/input1.wav test_wav/input2.wav test_wav/input3.wav

OpenAI-compatible Speech-to-Text API

UWebASR provides an OpenAI-compatible audio transcription endpoint, so it can be used as an alternative to OpenAI speech-to-text models such as whisper-1 and gpt-4o-transcribe in applications that support a custom API endpoint.

Use the following transcription endpoint:

https://uwebasr.zcu.cz/v1/audio/transcriptions

Set the STT model to any model ID available in the API model ID list, for example generic/cs/zipformer. The API does not require an API key. If a client requires a non-empty key, enter any placeholder value; UWebASR ignores it.

Example request

curl https://uwebasr.zcu.cz/v1/audio/transcriptions \
  -F model=generic/cs/zipformer \
  -F file=@audio.wav

The multipart fields must be sent in this order: model first and file second. UWebASR needs the model ID before it starts processing the streamed file body.

Request parameters follow the OpenAI Speech-to-Text API documentation.