Clips from a range of categories — facial expression, high body motion, multiple characters, animation and emerging objects — segmented from full-length video across many genres.
Sample
Fields on the left describe every attribute; the example on the right is real JSON from one asset in this dataset.
Each pair-clip is one matched span across two different camera recordings of the same real-world event. source_1 and source_2 point to the two source assets and their time ranges; framewise_sim reports the cosine-similarity band (mean / min / max) across the matched frames.