Data Marketplace

Use verified, licensed data with confidence. You can download right away or check the data through inquiry.

Sell your data on Flitto

Simply register your data and check its sale eligibility.

A total of 323 datasets
  • Pre-training DataAudio

    English-Vietnamese Parallel Speech Dataset

    A parallel speech dataset consisting of sentence-level aligned English and Vietnamese utterances, designed for training Speech-to-Speech translation models.

  • Pre-training DataAudio

    English-Indonesian Parallel Speech Dataset

    A parallel speech dataset consisting of sentence-level aligned English and Indonesian utterances, designed for training Speech-to-Speech translation models.

  • Pre-training DataAudio

    Japanese Natural Dialogue Speech Dataset

    A speech dataset containing natural two-party dialogues in Japanese. It captures everyday conversational flow, intonation, and turn-taking timing, suitable for training general-purpose speech recognition and dialogue models.

  • Pre-training DataAudio

    English-Korean Parallel Speech Dataset

    A parallel speech dataset consisting of sentence-level aligned English and Korean utterances, designed for training Speech-to-Speech translation models.

  • Pre-training DataVideo

    Physical AI: Human-Object Interaction Egocentric Dataset

    An egocentric human-object manipulation video dataset collected for training Physical AI models in everyday environments. Includes original footage, ECoT annotations, hand keypoint and skeleton visualization videos, and structured metadata such as trajectory data.

  • Frontier DataText

    Multilingual Chain-of-Thought Reasoning Text Dataset

    A multilingual chain-of-thought reasoning dataset built from complex problems requiring step-by-step decomposition and coherent answer generation, with AI-generated drafts reviewed by expert-level annotators.

  • Frontier DataText

    Expert CoT Text Dataset

    An expert chain-of-thought text dataset built from expert verbal reasoning to support LLM training for step-by-step reasoning.

  • Frontier DataText

    Doctoral Exam Questions and Solutions Text Dataset

    A high-difficulty text dataset built from doctoral-level exam questions and solutions to support LLM training for expert reasoning and problem solving.

  • Frontier DataText

    Domain-Specific Benchmark Dataset

    A multi-turn benchmark dataset built by benchmarking BFCL to evaluate agent action performance across finance, legal, medical, manufacturing, and defense domains.

  • Frontier DataText

    Safety Response Multi-turn Dataset

    A multi-turn conversational dataset designed to evaluate model response capabilities against major safety risk categories and attack patterns.