ukSOLARS audio data

Audio training data · already transcribed · word-level timestamps

Audio training data at $0.15 an hour, already transcribed.

50M+ hours of speech from public sources, across 90+ languages and a wide range of accents. Every hour comes with an AI transcript and word-level timestamps, so it is ready to train speech-to-text, text-to-speech and voice agent models as soon as you get it.

ac-0e91fEnglish·General American

nobodynoticeslatencyuntilthemomentitstopsbeingthereandthenit'salltheyhear

0.00/5.42s
Every word has a start time, an end time and a confidence score
Hours in catalog
50M+
Languages
90+
including code-switched speech
Transcript
Word-level
start and end time per word
Audio
48 kHz Opus
96–128 kbps
Use cases

What you can train with it.

Matched audio and text is what speech recognition, speech synthesis and multimodal models train on. The same hours work for all three.

Speech to text · ASR / STT

Train and test speech recognition

Conversational, accented and noisy audio, including phone calls — closer to what your users actually say than scripted studio recordings.

  • Word-level timestamps work directly with CTC, RNN-T and Whisper-style training
  • A confidence score on every word, so you can filter before training
  • Code-switched speech for multilingual models
  • Long recordings for streaming and segmentation tests
Text to speech · TTS

Train text-to-speech voices

Natural timing, pauses and emphasis from unscripted speech, with the alignment you need to model how long each word takes.

  • Word timings give duration targets and help with forced alignment
  • Many different speakers, for voice variety
  • Conversational and expressive speech, not just narration
  • Material for accent and cross-language transfer
Voice agents · speech LLMs

Pretrain and fine-tune multimodal models

Each hour is a matched pair of audio and text, which is what speech LLMs, multimodal large language models and voice agents train on.

  • Audio and text pairs for pretraining and fine-tuning
  • Real turn-taking and interruptions
  • Calls, interviews, lectures and media
  • Enough volume to pretrain, not just fine-tune
Sample data

See the data before you buy.

This is the manifest that comes with every order: one row per audio file, with its word-level timestamps as JSON. Ask for a sample and we’ll send audio and transcripts for the languages you need.

manifest.jsonl — 9 rows shown
audio filedurationratelanguageword-level timestamps
ac-0e91f.webm11:4248 kHzen-US[{"word":"nobody","start":0.18,"end":0.62},{"word":"notices","start":0.62,"end":1.08},{"word":"latency","start":1.11,"end":1.68},…]

Free 50-hour sample

We will send you 50 hours picked at random from the catalog, with their word-level transcripts, at no cost. No call needed.

Request sample data
Pricing

Estimate your cost.

$0.15 per audio hour at any volume, with transcripts included. Discounts start at 0.5M hours and increase from there.

100,000audio hours  ·  11.4 years of audio
Volume tiers
ListStandard rate, any volume$0.150
Volume0.5M hours and up$0.135
Scale1M hours and up$0.120
Corpus2.5M hours and up$0.105
Estimated total$15,000
Your rate
$0.150/hrList
Nearest competitor
$30,000
You save
$15,000

400,000 more hours unlocks the Volume tier at $0.135/hr.

Request this quote →

Estimated pricing. Your final rate is confirmed in your quote and depends on the language mix and any filtering you need.

Specifications

What you get.

Audio format
Opus in WebM, 96–128 kbps, 48 kHz, mono or stereo
Transcript format
JSON per file — word, start time, end time, confidence
Metadata
Language, duration, bitrate, audio file
Audio processing
Delivered as sourced, not re-encoded or normalised
Licence
Perpetual, worldwide, non-exclusive, for commercial model training
Delivery
Your S3, R2 or GCS bucket, or a direct download link — included in the price
Minimum order
10,000 hours
Payment
Upfront
Quote request

Request a quote.

Tell us what you need and we’ll reply with a price, how many hours are available in your languages, and a free sample.

Prefer email? trainingdata@uksolars.com

We only use this to answer you. No newsletter, and we don’t share it.

FAQ

Frequently asked questions.

Where does the audio come from?

Publicly available sources. The catalog is speech first — there is little to no music in it. You receive the audio files themselves, with language and duration metadata.

What does one hour include?

One hour of audio plus its transcript and metadata. Transcripts and word-level timestamps are included in the $0.15 per hour, and so is delivery.

How accurate are the transcripts?

They are AI transcripts with a confidence score on every word, so you can filter before training. Most are verbatim and include punctuation and capitalisation. Request the free sample and measure it against your own reference set.

What is in the free sample?

50 hours selected at random from the catalog, with their transcripts, at no cost.

What can I do with the data?

You get a perpetual, worldwide, non-exclusive licence to train commercial models on it. Redistributing the data itself is not permitted.

Is the data exclusive to me?

No, it is non-exclusive by default, which is what keeps the price at $0.15 per hour. Exclusive arrangements are available on larger orders.

Can I choose which languages I get?

Yes. Tell us the language mix you want and we will tell you how many hours are available.

How does payment work?

Upfront.