- The Spotify API offers dozens of audio and context variables (energy, valence, duration, danceability, etc.) that allow you to model and understand what makes a song popular.
- Statistical analysis shows that almost all audio features differ between popular and non-popular songs, while genre, tonality, or title sentiment have less predictive weight.
- Machine learning models such as Logistic Regression, KNN, SVM, Naive Bayes, and Random Forest achieve around 84–85% accuracy in classifying whether a topic will be in the top 15% most popular.
- The length of the track, its energy, loudness, and instrumentality stand out as key factors in its popularity on Spotify, in an ecosystem where editorial playlists and the recommendation algorithm are crucial.
Spotify has become the perfect laboratory to study what makes a song a hitWe have real-time playback data, advanced metrics for artists, and millions of listeners making decisions every second. Far from being just a listening platform, Spotify has transformed into a huge database where patterns, emotions, genres, and playlist strategies can be analyzed to understand why some songs take off and others sink into oblivion.
In recent years, tools, academic research, and machine learning models dedicated to Spotify hit analysisFrom studying editorial charts like Today's Top Hits or Rap Caviar, to building algorithms capable of predicting whether a track will belong to the most popular group in the catalog. At the same time, Spotify for Artists has democratized access to these metrics, allowing any artist to see who is listening to them, where and how they are being discovered, and adjust their creative and marketing strategy accordingly.
What exactly is Spotify hit analysis and why does it matter?
When we talk about Spotify hit analysis We're not just talking about looking at how many times a song has been played. It's a broader approach that combines audio, context, and user behavior to answer an uncomfortable but crucial question: can you predict a song's success before it's released? Using data from Spotify's API, researchers are working with variables like popularity, energy, valence, duration, tempo, and danceability to try to get closer to that answer.
This approach is rooted in the so-called “hit song science”This idea, popularized by Mike McCready in the early 2000s, involves using algorithms and mathematical models to predict which songs will perform well on charts and radio. While early studies concluded that this was very difficult (and that the models weren't accurate enough), the landscape has completely changed with the advent of streaming, the massive increase in data volume, and the refinement of machine learning techniques.
Today, Spotify offers an API with access to tens of thousands of tracks with its musical and contextual attributesFrom the key and mode, to the probability of a track being acoustic or instrumental, including metrics specifically designed to describe how the song is perceived (energy, valence, danceability, etc.). All of this allows us to move from subjective opinions to much more precise quantitative analyses.
In parallel, the tool Spotify for Artists It helps creators avoid getting lost in vanity metrics and focus on what truly matters: long-term audience development, engagement, retention, and how marketing efforts impact the creation of genuine fans. In other words, numbers, yes, but with context and strategic intent.
Official playlists and strategy: the weight of editorial lists
One of the most important pieces in any serious analysis of hits on Spotify is the role of the official editorial playlistsCharts like Discover Weekly, Release Radar, Today's Top Hits, New Music Friday, Rap Caviar, Mint or ¡Viva Latino! act as massive amplifiers: getting on one of them can completely change the trajectory of a song or even a career.
Third-party tools like Soundcharts allow you to see how an artist behaves in these playlistsNumber of appearances, length of stay, markets where they have the most impact, etc. This type of analysis makes it clear that presence in editorial lists is not an ornament, but a key component of the growth strategy on the platform.
There are certain recurring criteria in the construction of quality playlists. For example, it is valued that there are a variety of artists instead of repeating the same names over and over againLists where the same artist appears too many times tend to get worse ratings, because they reduce the listener's sense of discovery.
Also looking for a coherent gender balanceWhen a playlist mixes too many styles without a clear direction, the experience It becomes fragmented and the rating drops. The best-perceived lists tend to focus on one or a few well-defined genres, which helps the user understand at a glance what to expect when they press play.
Another key aspect is the a mix of well-known and emerging themesPlaylists that only include massive hits are fine, but they offer little in the way of discovery. In contrast, playlists that combine established songs with lesser-known gems tend to generate greater engagement and help create new hits.
Finally, there is some consensus that a good playlist should have at least around 50 tracks to offer a complete experienceCharts with fewer than 10 songs are often perceived as poor and receive lower scores, which also affects the performance of the tracks they include.
A brief history of Spotify and the origin of its analytics
Before Spotify could perform advanced hit analysis, the platform itself had to find its place in the music industry. The idea, conceived by co-founder Daniel Ek, arose after Napster's collapse in 2002. He wanted to create a service that was "better than piracy, but that also paid the industry."At that time, Ek was running uTorrent, one of the major P2P download and sharing clients, so he had firsthand knowledge of the world of unauthorized distribution.
After selling uTorrent to BitTorrent in late 2006, Ek focused entirely on building Spotify. The application was officially launched in 2008 in Sweden, after reaching licensing and equity participation agreements with major music corporationsSony Music Entertainment, Universal Music Group, and Warner Music Group. A year later, it expanded to the United Kingdom and, in 2011, landed in the United States.
In that initial period, the paid subscriber base grew from around 1 million in Europe to approximately 4 million globally in 2012By 2016, Spotify was already announcing 40 million paying users and around 100 million total users, consolidating streaming as the dominant form of music consumption worldwide.
At the same time, data tools for artists began to appear. First was Fan InsightsThis solution offered some teams limited access to streaming data: demographics, geography, and basic trends. In 2017, this solution evolved into Spotify for Artists, opening the door for all artists to see their key metrics and use the data to make decisions about touring, releases, and promotion.
In 2018, the company went public with a market capitalization of approximately 30.000 millionSince then, the number of markets where it operates has continued to grow, and with it, the importance of analytics within the platform. What began as a licensed media player has become a global infrastructure where data reigns supreme.
Spotify metrics: how to measure a song beyond streams
To build hit prediction models, it is first necessary to understand what type of data provided by the Spotify APIIn one study focused on the Indian market, for example, more than 46.000 song records were extracted from playlists generated by Spotify, covering a wide variety of genres and subgenres.
The information is organized into several sections. At track level, we find data such as identifier, title, artist, popularity, release date, and duration in millisecondsAt the album level, the ID and name are stored. At the playlist level, the name, ID, genre, and associated subgenre appear.
But the most interesting thing for hit analysis is the audio featuresdivided into different categories: “mood” characteristics (danceability, energy, valence, tempo), physical properties (loudness, speechiness, instrumentalness), context (acousticness, liveness), and musical segments (key, mode). Each of these variables is carefully defined and normalized.
Danceability, for example, is expressed as a value between 0 and 1 that summarizes how easy is it to dance to a songBased on elements such as rhythm, tempo stability, and beat strength, the energy also ranges from 0 to 1, attempting to capture whether the track sounds intense, fast, and powerful, versus something calmer or softer.
Valencia describes the perceived emotional “positivity” In audio, low values correspond to sad, tense, or somber feelings, while high values fit with cheerful, bright, or euphoric music. Acousticness measures the likelihood of a track being acoustic, liveness the presence of a live audience in the recording, and speechiness reflects the proportion of spoken words in the mix (very high, for example, in podcasts or spoken-word tracks).
Other important parameters are the tempo in BPMthe total duration, the key (encoded as an integer), the mode (major or minor), and the track's popularity. The latter is an internal Spotify metric that ranges from 0 to 100. It depends both on the volume of streams and how recent they areA song that was huge years ago, but is hardly played anymore, will see its rating drop over time.
How to prepare and clean the dataset to be able to predict hits
One step that is often overlooked when discussing Spotify hit analysis is the data set preparationIn the study at hand, the data were collected using R and RStudio, calling the Spotify API for different combinations of market (India), genres and subgenres selected according to their global and local weight.
When combining songs from various themed playlists, it's common for many tracks to be repeated, because the same song can appear in different lists. Since the objective wasn't to analyze the playlisting strategy itself, but rather the characteristics of the songs, it was They removed the duplicate tracksreducing the total base from 46.417 to approximately 39.147 unique songs.
Columns that were not going to be used as explanatory variables in the models, such as artist, album ID and name, and playlist-specific fields, were removed. At the same time, They standardized the field names and the data types were adjusted: for example, popularity, mode, key and duration were converted into float-type numeric values to facilitate statistical treatment.
A key decision was to transform popularity into a more manageable variable. Instead of working with a continuous value from 0 to 100, a threshold for separating popular and unpopular songsTaking the 85th percentile of the distribution (around a popularity of 65), tracks above that value were considered "popular" and the rest "not popular".
In this way, the dataset was divided into almost 6.000 popular songs and about 33.000 non-popular onesFor the exploratory analysis, the data was even segmented into five popularity classes (very high, high, medium, low, and very low), with intervals of 20 points, which allowed for the comparison of fine trends between levels of success.
In addition, a new variable was generated from the song titles: a sentiment indicatorUsing the TextBlob library in Python, the polarity of each title (a number between -1 and 1) was calculated and classified as positive, negative, or neutral. This numerical value was added to the dataframe to study whether the tone of the title is linked to its popularity.
Data exploration: what distinguishes the most popular songs
Before training any machine learning model, it is essential to dedicate time to the visual and statistical exploration of the dataIn the case of this study, distribution curves, averages per popularity group, and relationships between emotions and success were analyzed.
One of the striking findings relates to the popularity of the most successful songs. When looking at tracks with a popularity rating above 90, it was observed that A greater number of them were below the value of 0,5 in ValenciaIn other words, they sounded sadder, more somber, or melancholic than cheerful. It wasn't an absolute dominance, but it was a clear tendency.
When the distribution of popularity by gender was plotted, most curves adopted a approximate bell shapewith many tracks around the average and fewer at the extremes. However, genres like rock, R&B, EDM, world, and Indian music showed some asymmetry due to a large number of tracks with zero popularity. In contrast, styles like pop, rap, Latin, and desi appeared more balanced and with less concentration of zero-rating tracks.
This suggests that in those more mainstream genres it is, in general, easier to reach a certain level of popularityeven if the songs do not have all the ideal characteristics, while in other styles the distribution of success is more polarized.
When comparing the averages of the different features by popularity class, a very clear pattern emerged: the most successful topics tend to have greater energy and greater loudness (perceived volume)as well as more instrumentalism and less speechiness. Simply put, songs that perform best on streaming tend to be more powerful and more focused on music than on spoken words.
It was also observed that popular tracks feature lower acousticness and lower livenessIn other words, they sound less "acoustic" and less "live." They are more studio-produced, more polished, with less ambient noise or concert feel.
Another key finding is that popular songs tend to be shorter than the unpopular onesIn a context where revenue depends on the number of streams and algorithms reward repetition, it makes sense that shorter tracks can accumulate plays more quickly and become more "friendly" for playlists.
Regarding perceived happiness (valence), the trend is curious: levels rise from the lower popularity classes to the "upper" class, but in the most extreme category, the "very high," there is a marked drop. In other words, Super-popular songs tend to sound sadder than merely "successful" ones.This opens the door to many interpretations about public taste and the socio-emotional context.
Statistical analysis: impact of genres and audio attributes
Once the data had been explored, the next step was to apply formal statistical tests to see which variables could be considered relevant in explaining popularity. To study whether gender influenced success, an analysis of variance (ANOVA) was used with nine main genres.
The ANOVA result yielded a very high F-value and a practically zero significance, which means that There are statistically significant differences in popularity between genresHowever, that is not enough: it is also important to know between which pairs of genders these differences occur and whether the variable will be useful for a predictive model.
Since the variances were not homogeneous and the number of songs per genre varied, the following was used: Games-Howell post-hoc testThis analysis made it possible to verify which gender combinations did not show significant differences in popularity, which, overall, made it less advisable to use gender as a strong predictor in machine learning models.
In parallel, the following were carried out independent t-tests For each audio feature, the means were compared between the group of popular songs and the group of non-popular songs. With a significance level of 95%, it was found that almost all variables (danceability, energy, loudness, acousticness, liveness, duration, tempo, valence, and instrumentalness) showed significant differences between the two groups.
The only apparent exception was the speechinessIn the first test, the difference in means was not significant, but further analysis revealed that the speechiness ranges for the highly popular topics (with values between 0,024 and 0,685) fell within the broader range of the less popular class (between 0 and 0,964). This overlap did not prevent speechiness, when used effectively, from contributing information to the model, so it was decided to retain it as a predictor.
In summary, the audio features block demonstrated a clear descriptive power regarding popularity, while gender was handled more cautiously and was ruled out as a primary variable in some prediction approaches.
Building machine learning models to predict hits
With the clean dataset and the relevant variables selected, it was time to create binary classification models capable of predicting whether a song belongs to the top 15% of popularity. To do this, the categorical variables key and mode were first converted into dummy variables using Pandas' get_dummies function.
After incorporating these dummies, the original columns for key and mode, as well as continued popularity, were removed, and the target variable was retained. binary field of popularity (popular vs. not popular). Before training the models, all numerical features were standardized with StandardScaler, since they worked on very different scales and it was important that they had a comparable weight.
The dataset was split into training and test using the train_test_split function, reserving the 20% of the registrations for testingIn this way, information leaks were avoided and it was ensured that the performance metrics reflected the models' actual generalizability, and not just their ability to memorize training data.
The first model applied was the Logistic regressionThis technique is widely used in binary classification. An iterative version with a learning rate of 0,01 and 200 iterations was implemented and evaluated using cross-validation. The average accuracy reached approximately 84,7%, a remarkably strong result for a problem as complex as predicting musical success.
Then the algorithm was tested K-Nearest Neighbors (KNN)which classifies each song according to the k nearest neighbors in the feature space. Adjusting the value of ky using Grid Search to find the optimal configuration yielded very similar accuracy, around 84,8%, with a slight improvement in the cross-validation score compared to logistic regression.
The next model was the Support Vector Machine (SVM)Another supervised technique capable of handling both linear and nonlinear relationships was also used. Again, through hyperparameter fitting, accuracies of around 84,6% were achieved, with equivalent values in cross-validation, indicating stable performance.
The following was also evaluated: Naive Bayes classifierBased on the hypothesis of independence between features. Despite being a much simpler model in terms of assumptions, its accuracy results (approximately 84,6%) were on par with SVM and very close to logistic regression and KNN, reinforcing the idea that the problem structure lends itself well to this type of probabilistic approach.
Finally, models based on decision trees were tested. Decision Tree Classifier It achieved a significantly lower accuracy, around 75,4%, and cross-validation confirmed this drop in performance, likely due to overfitting issues. Random Forest Classifier, which combines many trees to smooth out these effects, improved significantly, achieving around 84,1% accuracy and over 84% in cross-validation.
To complete the analysis, a confusion matrix and a classification report for the Random Forest. The model showed an excellent ability to identify non-popular songs (majority class), with an accuracy of 0,85 and a recall close to 0,99, while its performance in the popular song class (minority) was more modest, with an accuracy of 0,40 and a recall of 0,05. This highlights the classic challenge of class imbalance in this type of problem.
Importance of variables: what weighs most when creating a hit
Beyond knowing which model is more accurate, it is key to understand which features actually provide useful information when predicting whether a song will be popular. For this purpose, XGBoost (XGBClassifier) was used to obtain variable importance scores, based on the contribution of each feature to error reduction in the decision trees.
The results showed that the The song's length clearly stood out from the restwith an F score well above 400. This finding fits with the idea that shorter tracks encourage repeated listening and therefore accumulate more plays in less time, something that Spotify's recommendation algorithm tends to reward.
At the opposite extreme, variables such as the key, the mode, or the feeling taken from the title They obtained importance scores below 50, suggesting that their contribution to the overall predictive power of the model was small. This does not mean they have no effect at all, but rather that, compared to other features, their weight is much less.
The remaining audio attributes (such as energy, danceability, valency, loudness, acousticness, liveness, speechiness, and instrumentalness) were in an intermediate range with F-scores between 200 and 300, indicating that They form a fairly balanced information blockNone of them completely dominates, but all together they help to build a fairly accurate picture of a song's chances of success.
All told, the models based on audio features were able to predict with around an 84-85% success rate if a clue would be part of the most popular 15%This lends some empirical support to the old intuition of "hit song science": with good data and appropriate techniques, musical success is not purely random, although there remains a significant unpredictable component.
The picture painted by this analysis shows that combining Spotify data, traditional statistics, and machine learning models allows us to go far beyond simply counting streams. Understanding the impact of duration, energy, valence, and instrumentality, along with the role of editorial playlists and the platform's recommendation dynamics, provides artists, labels, and analysts with a very concrete roadmap for interpreting why certain songs become hits and how they can maximize their chances of joining that select group without completely losing the human and creative element that, fortunately, remains irreducible to any formula.