Clarification on PAN@CLEF 2026 Generated Plagiarism Detection Leaderboard and Final Ranking

Hello,

I have a question regarding the PAN@CLEF 2026 Generated Plagiarism Detection task leaderboard on TIRA.

Could you please clarify whether the current leaderboard (TIRA) represents the final official ranking, or if it is still provisional? If it is not final, could you also specify which metric(s) will be used to determine the final ranking (e.g., nDCG@10, Recall@10, or a combined score)?

Additionally, when should we expect the official final results to be published?

Thank you very much for your time and clarification.

Thanks for reaching out!

The current table is our evaluation that we have at the moment. The overview paper that we will write after the notebook papers have been submitted will likely contain a complete evaluation and interpretation. There we will clarify and interpret the evaluations and experiments (we need to incorporate the notebook papers to do that properly).

It might also be that other/additional factors will be included into the evaluation. We need time to finish this, so the current evaluation table linked above is a preliminary one, it might be that we modify it, but it also might be that it remains similar or the same.

The final evaluation results will be published and discussed at the conference, everything before that is preliminary, but we do our best that the evaluations are already meaningful.
I personally do not see the experiments and evaluations as ā€œleaderboards/rankingsā€, usually I rather like to see it as different evaluation measures capture different user behaviors. In the shared task we collaboratively explore the space of potential solutions. There is no single approach that works always the best, everything has advantages and disadvantages. I think Recall@10 and nDCG@10 are important evaluation measures for the task. (but other measures also can be important, for instance when users only want to very fast decide if an paper should go into a deeper analysis or not, reciprocal rank might capture this user behavior.)

In the notebook paper for participating approaches, one can also build an own interpretation of the evaluation (or one can skip the evaluation part, it is enough when the participation notebook describes the approach).

Best regards,

Maik

1 Like