Hi everyone,
I’m with Sonexis. We’re currently expanding our supplier network for Indic-language speech and conversational data.
I’m looking to connect with organisations or individuals who own datasets or have documented authority to license them commercially.
We’re especially interested in existing data around:
- regional and accented speech
- multilingual / code-switched conversations
- ASR training and evaluation
- TTS
- telephony and call-centre speech
- voice-agent evaluation
- spontaneous and multi-speaker conversations
We care about more than total hours.
For us, the important questions are: where did the data come from, who can license it, what consent exists, what metadata comes with it, and what the dataset is actually useful for.
If you have something relevant, feel free to DM me.
You can also reach us at [partner@sonexis.in](mailto:partner@sonexis.in) or apply here: https://sonexis.in/suppliers/apply
If there looks to be a genuine fit, we can set up a call and go through the dataset properly.
Even if you’re not sure whether your data fits, feel free to send the basics: language, data type, approximate volume, collection method and rights position
submitted by /u/Cautious-Today1710
[link] [comments]