On the Human and Computer Alignment of Attribute-Based Music Matches
A perceptual experiment on music matches was conducted, focusing on five musical attributes: melody, harmony, rhythm, voice, and timbre. The study used a triplet-based forced-choice task with 300 cases, including plagiarism examples, cover songs, and AI-generated music. The resulting MATCHA dataset contains 1105 perceptual assessments from 83 expert participants. Findings show measurable agreement among participants in identifying matches across attributes and partial alignment between human judgments and computational approaches.
Recent advances in generative AI raise ethical concerns about originality and potential replication of training data, with implications for transparency, attribution, and intellectual property. In music, computational approaches using audio-based similarity metrics have been proposed to identify potential replication, but their alignment with human judgments across distinct musical attributes remains underexplored. To address this gap, researchers conducted a perceptual experiment on music matches, defined as strongly similar musical excerpts, focusing on melody, harmony, rhythm, voice, and timbre. The study employed a triplet-based forced-choice task with 300 cases, including plagiarism examples, cover songs, and AI-generated music. From this experiment, the MATCHA (Musical Attribute-based Triplet Comparison with Human Annotations) dataset was introduced, comprising 1105 perceptual assessments from 83 expert participants. Findings reveal measurable agreement among participants in identifying matches across attributes and partial alignment between human judgments and computational approaches.
The study uses a triplet-based forced-choice paradigm to collect human similarity judgments across five musical attributes. The MATCHA dataset enables quantitative comparison between human perception and audio-based similarity metrics. Partial alignment suggests that current computational metrics capture some but not all aspects of human-perceived musical similarity, indicating room for improvement in attribute-specific modeling.
This research addresses a critical gap in AI-generated music: the lack of reliable, human-aligned methods for detecting potential replication or plagiarism. As generative music tools proliferate, platforms and rights holders need robust similarity metrics that reflect human judgments to manage copyright and attribution concerns. The MATCHA dataset could serve as a benchmark for developing such metrics.
For music streaming platforms, record labels, and AI music startups, a human-aligned similarity metric could reduce legal risks and improve content moderation. The MATCHA dataset provides a foundation for building such tools, potentially creating new products or features for copyright detection and originality assessment.
Next signals to watch include: adoption of MATCHA as a benchmark in music information retrieval research; development of new audio similarity metrics trained or evaluated on MATCHA; and potential integration of human-aligned similarity measures into content ID systems or AI music generation tools to flag potential replication.