RecSys Challenge 2026

About top

The RecSys 2026 Challenge will be organized by Seungheon Doh (Korea Advanced Institute of Science and Technology, South Korea), Sergio Oramas (Pandora/SiriusXM), Bruno Sguerra (Deezer Research), Abhinav Bohra (Amazon), Claudio Pomo (Politecnico di Bari, Italy), and Francesco Barile (Maastricht University, Netherlands).

The RecSys Challenge 2026: Music-CRS focuses on the evolving landscape of music discovery, where static recommendation lists are being replaced by dynamic, conversational interactions. As users increasingly interact with AI through natural language, there is a critical need for systems that can seamlessly integrate Natural Language Understanding (NLU) with high-precision Recommender Systems (RecSys). This challenge aims to push the boundaries of how AI understands nuanced user preferences, explores musical tastes through dialogue, and provides contextually relevant track recommendations.

By utilizing the TalkPlayData-Challenge dataset, a large-scale conversation resource generated through an advanced agentic pipeline, we invite the global research community to tackle the complexities of multi-turn preference elicitation. As a research community-driven initiative, the challenge dataset features LLM-generated multi-turn dialogues paired with music metadata and user-item interaction data derived from publicly available research datasets. The dataset does not include proprietary or confidential data provided directly by SiriusXM, Deezer, or any challenge sponsor. This challenge serves as a bridge between the NLP and RecSys communities, fostering next-generation innovation in interactive and personalized music information retrieval.


Challenge Task top

The primary goal is to develop a Conversational Music Recommendation system that acts as an intelligent agent capable of navigating user tastes through dialogue.

Main Task: Conversational Music Recommendation. The system must understand user music preferences from previous conversation turns and user profiles to recommend relevant tracks from a catalog while generating natural, helpful responses.

Candidate catalog rule. During inference, recommender systems must retrieve candidates from the entire track catalog. Participants must not filter, subset, or restrict tracks using track_split_types or any other mechanism. For BM25/BERT baselines, the configuration must include track_split_types: ["all_tracks"]; submissions that do not use all_tracks may be considered invalid.


Evaluation top

The challenge employs a multi-dimensional evaluation framework to assess both what a system recommends and how it communicates those recommendations. The official composite score is:

Score = 0.50 x nDCG@20 + 0.10 x Catalog Diversity + 0.10 x Lexical Diversity + 0.30 x LLM-as-a-Judge

Development vs. blind evaluation. The public evaluator supports transparent development-set evaluation. Blind A served as an interim leaderboard phase, while Blind B is the official final evaluation split. The final ranking is based on the Blind B leaderboard.

Dimension Weight What it measures How it is computed Role in evaluation
nDCG@20 0.50 Ranking quality of the recommended tracks. Computed from the ranked list of predicted tracks against the ground-truth relevant item. Higher-ranked correct recommendations receive more credit. Primary recommendation metric.
Catalog Diversity 0.10 How broadly a system covers the music catalog. Number of unique recommended tracks across all predictions divided by the total catalog size. Complementary diversity indicator.
Lexical Diversity 0.10 How varied the generated language is. Measured with Distinct-2, i.e., unique bigrams divided by total bigrams across generated responses. Complementary response-generation indicator.
LLM-as-a-Judge 0.30 Quality of the generated explanation. Blind-set responses are judged by a Gemini model used as an automatic judge. The judge evaluates two text-only dimensions: Personalization and Explanation Quality. These dimensions evaluate the written response independently from recommendation accuracy. To preserve the integrity of the blind evaluation, we disclose the judge family but do not publish the evaluation prompt. Blind-set response-quality evaluation.

Aggregation policy. Each conversation turn has exactly one ground-truth track. The evaluation reports nDCG@1, nDCG@10, and nDCG@20, with nDCG@20 as the primary retrieval metric. Results are macro-averaged across sessions and turns. LLM judge scores use a 1-5 integer scale and are normalized to [0, 1] before being weighted in the composite score.

To preserve the integrity of the blind evaluation, the judge family is disclosed, but the detailed evaluation prompt is not published. Blind-set scoring is performed on the official Codabench leaderboard infrastructure.


Dataset: TalkPlayData-Challenge top

The challenge uses TalkPlayData-Challenge, a large-scale multi-turn dialogue dataset for conversational music recommendation, with pre-extracted multimodal track and user embeddings provided. It is based on the publicly available TalkPlayData dataset and was prepared for this research challenge by members of the organizing committee affiliated with KAIST.

The leaderboard evaluation proceeds in two blind stages. Blind A supported the interim leaderboard during the main phase of the challenge. Blind B is the final hidden evaluation set and the final leaderboard is based on Blind B. The dataset is released under CC BY-NC 4.0 for non-commercial research use only; redistribution outside the scope of the challenge is not permitted.

Dataset Components

For further details, please refer to the dedicated website .


Prize top

The total prize pool is $2,000 USD, funded by external sponsors including SiriusXM and Deezer. Sponsors provide financial support only and do not control or influence the dataset, evaluation methodology, rankings, or outcomes of the challenge.

Winners are responsible for any applicable taxes in their jurisdiction.



Winners top

Blind-Dataset-B (Final Leaderboard)

All · 40 teams · Composite Score, descending

Final

Composite Score = 0.50 × nDCG@20 + 0.10 × Catalog Diversity + 0.10 × Lexical Diversity + 0.30 × LLM-as-a-Judge. LLM-as-a-Judge scores are normalized before aggregation.

# Team ID Type Composite
Score
nDCG@20 Catalog
Diversity
Lexical
Diversity
LLM-as-a-
Judge
Code
Loading leaderboard...

Scroll horizontally to see all metrics on smaller screens.

* Additional verification is currently in progress. Source: official Music-CRS Challenge results.



Timeline top

The dates below follow the official Codabench Timeline page. Dates may change; participants should refer to Codabench for the latest schedule.

When? What?
31 March, 2026 Website published.
10 April, 2026 Start of the RecSys Challenge; Train, Development, and Blind A datasets released.
17 April, 2026 Submission system opens; leaderboard live for Blind A.
23 June, 2026 Blind Dataset B released; Blind B final phase opens.
30 June, 2026 End of the RecSys Challenge.
6 July, 2026 Final leaderboard and winners announced; EasyChair opens for paper submissions.
9 July, 2026 Code upload deadline for final predictions.
20 July 24 July, 2026 Paper submission deadline.
3 August 5 August, 2026 Paper acceptance notifications.
10 August, 2026 Camera-ready papers due.
September 2026 2 October, 2026 RecSys Challenge Workshop at ACM RecSys 2026.

RecSys Challenge Workshop

Workshop Program and Accepted Papers

Talks, research papers and breaks throughout the workshop day.

Back to top ↑
Friday, 2 October 2026 All times are local
  1. 8:30–8:35
    Opening Remarks
  2. 8:35–9:25
    Keynote 140 min + 10 min Q&A Enrico Palumbo Spotify From Language to Items: Building Conversational Recommendation Experiences at Spotify
  3. 9:25–9:40
    Challenge Presentation: Analysis and Insights
  4. 9:40–10:00
    Paper Session 1: Auditing the benchmark (labels vs requests, metric consistency)
  5. 9:40–9:50
    Auditing Alignment among Item Relevance, Goal Progress, and LLM-Judged Responses in Music-CRS (Paper 1280, Industry) — Tomoya Terai
  6. 9:50–10:00
    When the Label Ignores the Request: Auditing Policy-Selected Targets in Synthetic Conversational Music Recommendation (Paper 1286, Industry) — Sanjeev Suresh
  7. 10:00–10:30
    Coffee Break
  8. 10:30–12:00
    Paper Session 2: Building the system, from multi-source retrieval to grounded responses
  9. 10:30–10:40
    Talk Less, Rank Better: A Decoupled Pipeline for Conversational Music Recommendation (Paper 1289, Academic) — Abdallah Alkhetiar, Luigi Inguaggiato, Nicolò Locatelli, Roberto Manea, Alessio Pizzini, David Ravelli, Gianmarco Schifone, Matteo Vitali, Andrea Zhang, Michael Benigni, Andrea Pisani and Maurizio Ferrari Dacrema
  10. 10:40–10:50
    A Practical Multi-Source Pipeline for Conversational Music Recommendation in the RecSys Challenge 2026 (Paper 1270, Industry) — Ryohei Wakatsuki
  11. 10:50–11:00
    A Multi-Source Retrieve–Fuse–Rerank System for Conversational Music Recommendation (Paper 1282, Academic) — Youness Soussou, Loubna Mekouar and Youssef Iraqi
  12. 11:00–11:10
    Dialogue-Aware Conversational Music Recommendation via Multi-Modal Retrieval and Weighted RRF (Paper 1272, Academic) — Aditya Rai, Sanskar Aggarwal, Venkatesh Shukla and Amandeep Kaur
  13. 11:10–11:20
    Dialogue-Aware Music Recommendation via Fused Retrieval and Learned Ranking for the TalkPlayData Challenge (Paper 1276, Industry) — Suryaa Veerabathiran Seran
  14. 11:20–11:30
    MiniMaestro: Resource-Conscious Conversational Music Recommendation with a Single Open-Weight 8B Model (Paper 1287, Industry) — Simran Sundrani and Mohan Bhambhani
  15. 11:30–11:40
    Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation (Paper 1274, Industry) — Sungwook Yoo, Sewook Yoo
  16. 11:40–11:50
    Picking is Not Ranking, and Explanation Quality Has Many Dimensions: Lessons for Conversational Music Recommendation (Paper 1285, Academic) — Maxime Manderlier and Fabian Lecron
  17. 11:50–12:00
    Beyond Single-Signal Retrieval: A Graph-Aware Agentic Framework for Conversational Music Recommendation (Paper 1187, Academic) — Marco Valentini, Bianca Di Bitetto, Gianmichele De Palma, Berardino Como, Roberta Russo, Francesco Maria Dicataldo, Francesco Salvatore Desiderato, Francesco Falcone, Chiara Mallamaci, Antonio Ferrara, Daniele Malitesta, Tommaso Di Noia and Fedelucio Narducci
  18. 12:00–13:30
    Lunch Break
  19. 13:30–14:10
    Keynote 2 Domonkos Tikk Taboola Netflix Prize revisited: a hard problem even after 20 years
  20. 14:10–15:00
    Paper Session 3: Ceilings, shift and negative results (why gains do not transfer)
  21. 14:10–14:20
    Distribution-Robust Reranking for Conversational Music Recommendation: PoliBaJukebox at the ACM RecSys Challenge 2026 (Paper 1273, Academic) — Andrea Lops, Nicola Cipriani, Gabriele Colapinto, Marcantonio de Candia, Vito Di Bari, Mauro Foglia, Gledjan Meta, Davide Savoia, Antonio Ferrara, Daniele Malitesta, Tommaso Di Noia and Fedelucio Narducci
  22. 14:20–14:30
    Frozen Embeddings, LLM Retrieval Lanes, Critic-Guided Replies: Conversational Music Recommendation at an Information-Limited Ceiling (Paper 1281, Industry) — Artem Volgin, Ben-Orli Nathanson, Daniele Dal Grande, Alisa Milchevskaya, Adham Al Hossary and Rami Albatal
  23. 14:30–14:40
    Routing Discovery in Conversational Music Recommendation (Paper 1278, Academic) — Madhav Patil, Shreyas Biradar and Divyansh Bhatia
  24. 14:40–14:50
    An Empirical Analysis of Retrieval and Evaluation in the RecSys Challenge 2026 Music-CRS Track (Paper 1279, Academic) — Siwon Lee
  25. 14:50–15:00
    State-Driven Retrieval and Learned Re-Ranking for Conversational Music Recommendation (Paper 1275, Industry) — Nidhin Pattaniyil, Semih Yagli and Tanwir Zaman
  26. 15:00–15:30
    Coffee Break
  27. 15:30–16:30
    Poster Session
  28. 16:30–17:00
    🏆 Award Ceremony and Closing

Keynote Speakers top

Portrait of Enrico Palumbo

Enrico Palumbo

Enrico Palumbo is a Senior Research Scientist at Spotify, previously at Amazon Alexa. His research focuses on improving Search and Recommendations through Generative AI, with a recent interest in agentic systems and generative recommendation. He has been a core contributor to the design and launch of AI products used by hundreds of millions of users, including Spotify's Agentic Search, Query Autocomplete, and Conversational Agent, and Alexa's non-English models. His work has resulted in several patents and publications in top-tier venues such as RecSys, KDD, WebConf, CIKM, and ESWA. He holds a PhD in Knowledge Graph Embeddings for Recommender Systems, which he carried out jointly between the Polytechnic University of Turin, EURECOM, and Links Foundation.


Portrait of Istvan Pilaszy

Istvan Pilaszy

Istvan Pilaszy is a co-founder of Gravity R&D, a recommender systems company acquired by Taboola in 2022. He was a member of Gravity’s team in the Netflix Prize competition, which eventually joined The Ensemble, tying for first place in the final competition. He holds a PhD in Computer Science from the Budapest University of Technology and Economics, where his research focused on factorization-based large-scale recommendation algorithms. At Gravity R&D, Istvan was responsible for the core recommendation engine, ensuring its scalability, efficiency, code quality, and reliability. He developed new recommendation algorithms, brought them into production, and scaled them to real-world workloads. Gravity’s recommendation engine lives on at Taboola and continues to evolve. Since the acquisition, Istvan has focused on real-time bidding and adapting the engine to the needs of a much broader range of use cases.


Portrait of Domonkos Tikk

Domonkos Tikk

Domonkos Tikk is a co-founder of Gravity R&D, a recommender systems company acquired by Taboola in 2022, and currently leads Taboola’s R&D site in Hungary. He holds a PhD in computer science from the Budapest University of Technology and Economics. His research interests span recommender systems, machine learning, and text mining. He led Gravity’s team in the Netflix Prize competition, which eventually joined The Ensemble, tying for first place in the final competition. Domonkos has been an active member of the ACM RecSys community for more than 15 years, serving in numerous organizational roles, including RecSys Challenge Co-Chair, Industry Co-Chair, and Program Co-Chair of RecSys 2019. He is also a member of the ACM RecSys Steering Committee. He has co-authored approximately 200 publications, which have received more than 13,000 citations. He is also a recipient of an Alexander von Humboldt Research Fellowship for experienced researchers.


Paper Submission Guidelines top

Submission website: EasyChair

Important dates: Paper submission deadline: 20 July 2026; acceptance notifications: 3 August 2026; camera-ready deadline: 10 August 2026.

Important note to authors about ACM's new open access publishing model

ACM has introduced a new open access publishing model for the International Conference Proceedings Series (ICPS). Authors based at institutions that are not yet part of the ACM Open program and do not qualify for a full geographic waiver will be required to pay an article processing charge (APC) to publish their ICPS article in the ACM Digital Library. To determine whether or not an APC will be applicable to your article, please follow the detailed guidance here: ACM ICPS author guidance .

Further information may be found on the ACM website: full details of the new ICPS publishing model and full details of the ACM Open program .

Please direct all questions about the new model to icps-info@acm.org.

Terms & Conditions top

These terms summarize the official Codabench Terms & Conditions for the Music-CRS Challenge 2026. By registering, participants agree to comply with the official rules.

Official reference: Codabench RecSys Challenge 2026 .

Organization top

Organizing Committee top