Trained listeners misjudge AI and human music in blind tests
The study found that authorship guesses were weak when labels were removed, yet those guesses still shaped how people heard the music.
Musically trained listeners could not reliably tell whether short musical excerpts were AI-generated, human-AI collaborative or fully human-composed when they heard them without labels. Their beliefs about who made the music still affected the aesthetic scores they gave the pieces.
As reported on Aug. 21, SubmitHub said AI was involved in nearly 40% of July releases. In a Frontiers in Psychology paper published on Aug. 26, researchers tested blind-listening judgments of three kinds of music and asked whether listeners would still form opinions about authorship and quality without being told the source.
The experiment used 18 melodic excerpts produced through what the paper calls perceptual rule transformation, a process that turned compositional experience into executable sampling constraints through natural-language interaction. The system was built as a two-level collaborative framework with OpenAI Codex, combining a small Transformer that generated probabilistic melodic patterns with a rule-based layer that filtered outputs for musical conditions such as cadence, rhythm and tonality.
The researchers then asked 71 participants with music training to make creator-attribution judgments, rate their confidence and evaluate the music aesthetically. They also framed the work against a wider AI-music literature that, according to the paper, has focused heavily on technical implementation, creative output, creative agency and ethical or legal issues, while including comparatively few studies of composers’ own practices or the tool-development stage of AI tools.
The link between the actual compositional condition and attribution judgments was statistically significant but very small, with χ2(4)=23.1, p<0.001 and Cramér’s V=0.095. Judgments for the AI-generated and human-AI collaborative excerpts were close to random, and even in the human-composed condition 54.7% of excerpts were misattributed.
Aesthetic evaluation differed significantly across the three compositional conditions. After controlling for the actual condition, attribution confidence and individual differences, creator-attribution judgment remained independently associated with aesthetic evaluation, including within the AI-generated condition where the source of every excerpt was the same.
The paper concludes that, in blind listening, musically trained listeners could not reliably identify the compositional source of the excerpts, although they showed a modest advantage when the music was actually human-composed. More broadly, it suggests that listeners spontaneously infer a creative agent even when authorship labels are absent, and that those inferences can become a stable bias in how AI-generated music is received.