AI & ML interests

RLHF, Model Evaluation, Benchmarks, Data Labeling, Human Feedback, Computer Vision, Image Generation, Video Generation, LLMs, Translations

Recent Activity

Articles

Rapidata 's collections 8

Rich Human Feedback
1.5M responses from 152,684 people.
Video Model Releases
One dataset per launch, text-to-video and image-to-video. Veo, Sora, Kling, Pika, Seedance, ranked by real humans.
Research & Human Behaviour
What people think when no model is involved: where the future sits in your head, which sounds feel round, how our crowd compares to Prolific.
Benchmark Release Data
The vote data behind every live benchmark on benchmark.ai. One dataset per benchmark.
Text-to-Image Model Releases
One dataset per launch. A new image model ships, we put it in front of thousands of people, and publish how it actually did against other SOTA models.
Foundational Preference Sets
The 2M+ annotations our first benchmarks were built on: Flux, SD3, Midjourney and DALL¡E 3, split into preference, coherence and prompt alignment.
Source Sets & Smaller Studies
The raw material and the one-off runs including prompt sets, labeled image collections, and smaller studies.
Benchmark Release Data
The vote data behind every live benchmark on benchmark.ai. One dataset per benchmark.
Rich Human Feedback
1.5M responses from 152,684 people.
Text-to-Image Model Releases
One dataset per launch. A new image model ships, we put it in front of thousands of people, and publish how it actually did against other SOTA models.
Video Model Releases
One dataset per launch, text-to-video and image-to-video. Veo, Sora, Kling, Pika, Seedance, ranked by real humans.
Foundational Preference Sets
The 2M+ annotations our first benchmarks were built on: Flux, SD3, Midjourney and DALL¡E 3, split into preference, coherence and prompt alignment.
Research & Human Behaviour
What people think when no model is involved: where the future sits in your head, which sounds feel round, how our crowd compares to Prolific.
Source Sets & Smaller Studies
The raw material and the one-off runs including prompt sets, labeled image collections, and smaller studies.