|
|
Cyanite's Business Newsletter #15 |
|
Can you trust an LLM to analyze your music? |
LLMs can now describe music surprisingly well. But professional music workflows need more than a convincing answer. 👉🏼 Haven’t tried Auto-Tagging 2.0 with your own music yet? We’ve got something waiting for you below. |
|
If LLMs can process audio, why do we still need music tagging? |
In our last newsletter, we looked at why AI-powered workflows need reliable music understanding.But there’s an obvious question that follows:If today’s LLMs can already process audio, why do we need dedicated music analysis at all? |
|
Our Senior Research Engineer Benno Weck co-authored two benchmarks, MuChoMusic and HumMusQA, evaluating how audio-language models understand music.Together with two other recent studies, they form the basis of our latest article exploring where LLMs already perform well and where professional catalog workflows demand something more. |
|
|
For catalog workflows, consistency is key. |
Audio-capable LLMs can identify musical characteristics surprisingly well. But at catalog scale, consistency becomes critical.HumMusQA asked the same 320 music questions four times, changing only the order of the answer options. One model gave the same answer across all four runs for just 35% of the questions. |
|
This shows that even a small change in how a question is presented can affect results. For a single track, that variation may not matter much. For a repeatable catalog workflow across thousands or millions of tracks, it does. |
|
Is the answer grounded in the music? |
Consistency is only one part of reliable music analysis.Because LLMs predict the most likely answer based on the information available to them, their responses depend on the full context they receive. So another question becomes important: how much of an answer is actually coming from the music? To test this, researchers replaced audio with noise while keeping the same questions and answer options.If the model was truly using the music to find the right answer, removing it should lead to a big drop in performance.The result? One model still achieved around 38% accuracy, compared with around 56% when using the real recordings. |
|
Dedicated music tagging works differently. Cyanite’s models receive the audio signal as their input and predict defined musical attributes directly from it.With Auto-Tagging 2.0, we’ve expanded this approach to provide richer, structured music understanding for both human and AI-powered workflows. |
|
|
So, should you use an LLM to analyze your music? |
It depends on the task.Audio-capable LLMs are already useful for describing tracks and answering questions about music. But catalog workflows introduce additional requirements: results need to be consistent, grounded in the audio, and reliable at scale.That’s why we don’t see this as a choice between LLMs and dedicated music analysis. They can play different roles in the same system.LLMs and agents can enable incredibly flexible workflows. Reliable music understanding makes them robust enough at catalog scale. Together, they can enable powerful AI workflows across entire music catalogs. |
|
Give Auto-Tagging 2.0 a try! |
Many of you have been following our journey with Auto-Tagging 2.0 closely. As a thank you, we’d like to give you the chance to experience it firsthand and analyze a few tracks yourself! Give it a try and let us know what you think. We’d love to hear about your experience. |
|
|
Greetings and until next time, The Cyanite team! 👀 P.S. If you'd like to know what we've been up to (besides launching Auto-Tagging 2.0), check out the recap of our first-ever team offsite! |
|