AI Training Data

50,000+ Hours of Real B2B Call Audio — Dual-Channel, Outcome-Linked

Naturalistic enterprise conversations across collections, insurance, higher ed, and auto lending. Available for AI training data licensing.

Request a Data Card →
50K+
hours of audio
700K+
calls per month
112K
human annotations
95K
outcome-labeled pairs
9,500+
unique speakers
25%
bilingual ES/EN

The Dataset

Why this data is different

Most conversational audio datasets are scripted, acted, or scraped from podcasts. This is live, unscripted, high-stakes enterprise B2B — people negotiating debt payments, insurance claims, and enrollment decisions in real time. The emotional range, adversarial dynamics, and verified outcome labels are impossible to replicate synthetically. The longitudinal speaker tracking (9,500+ speakers over 2+ years) is rare in any commercial audio dataset.

Who This Is For

AI / ML Labs
Fine-tuning speech, audio LLMs, or conversational models that need real adversarial dialogue at scale.
Voice AI Companies
Building enterprise voice agents that need grounding in real objection-handling, interruption, and turn-taking patterns.
ASR / TTS Companies
Improving model accuracy on real-world conversational audio — noisy, overlapping, bilingual, and domain-specific.
Alt Data Buyers
Hedge funds and quant teams extracting consumer stress, payment behavior, and financial sentiment signals from live call data.
Sales AI Platforms
Ground-truth objection detection and recovery data for training coaching, scoring, and prediction models.
Request a Data Card

We'll follow up within 48 hours with format specs, schema documentation, and licensing terms.

Something went wrong — please try again or email anshul@altorlab.com

Thanks — we'll be in touch within 48 hours.

Check your inbox for a confirmation. In the meantime, feel free to reply with any specific questions about format, schema, or licensing.

Frequently Asked Questions

What verticals does the dataset cover?
Collections, auto lending, insurance, higher education, and energy. All calls are real enterprise B2B conversations — reps negotiating with customers on live calls, not scripted scenarios.
Is the audio dual-channel or mixed?
Dual-channel throughout. The rep and customer are on completely separate audio tracks with millisecond-accurate speaker separation — suitable for ASR fine-tuning, diarization training, and turn-level annotation tasks that single-channel audio cannot support.
What annotation format is used?
Annotations are at the utterance level, produced by domain experts. Labels include coaching markers, objection type, recovery strategy, and verified outcome per call (payment committed, closed, refused, escalated). Full transcripts and outcome labels are included per call.
Is PII redacted?
PII handling, consent framework, and redaction specifics are covered in the data card provided to qualified buyers after an initial inquiry.
What license terms apply?
Licensing is negotiated per engagement based on use case, scope, and exclusivity requirements. Common structures include research licenses, fine-tuning licenses, and commercial deployment licenses.
Is bilingual Spanish-English audio included?
Yes — 25% of the dataset is bilingual Spanish/English with naturalistic code-switching, as occurs in real US market calls. This is rare in existing commercial conversational audio datasets.