Hi UniAgent organizers,
We are currently validating our CIKM 2026 Task 2 (Solving) submission and would like to clarify the intended hosted LLM runtime configuration in TIRA.
Could you please confirm the following?
-
Will OPENAI_BASE_URL be provided or injected by the official TIRA evaluation environment?
-
Will OPENAI_MODEL also be provided by the evaluation environment, or should participants configure the model identifier themselves?
-
For the official Task 2 evaluation, should participants avoid supplying their own external API endpoint or API key?
-
Is there currently an official way to test the hosted-model execution path before final submission, or is this only available when running the code submission on TIRA?
-
Will the same submitted Docker image be evaluated against multiple hosted LLM endpoints/models without participants changing the configuration?
We want to make sure our local validation matches the official evaluation environment and avoid making assumptions about endpoint or model configuration.
Thank you!
Dear Hui,
Thanks for reaching out and for your participation in the task!
The overall idea is: Software should be configured so that it uses OPENAI_BASE_URL, OPENAI_MODEL, and OPENAI_API_KEY from its environment variables. We will then (after a software is uploaded) run it against three LLMs that we host internally. The LLMs that we will use during test are openai/gpt-oss-20b, openai/gpt-oss-120b, and Qwen/Qwen2.5-7B-Instruct all hosted via VLLM on our infrastructure. (We might use some additional models, but those models are our current plan, if you think we should include some other LLM, please do not hesitate to contact us, we could then maybe include this in our list.)
For the 5 questions:
1.) We will run all software that is submitted and we will configure the OPENAI_BASE_URL for the executions. You can use a LLM of your choice during development, but during test we need to run the software against LLMs that are only hosted in our private infrastructure as the data is confidential. In case we have questions on how to run the software, we will reach out.
2.) We will also provide the OPENAI_MODEL environment as well.
3.) Using an own external endpoint or API key is not possible, because the data is confidential and can not be send to a third party.
4.) We only have the four spot-check datasets that you can use on your end to develop your approach. This datasets are rather small, as it is the first iteration of the task we used all of our budget to create the test dataset. The evaluation results on the test data is blinded until after the deadline. In case a software fails, we unblind the execution to help fixing the potential problem. (We try to derive larger validation data that we maybe can synthetically generate from the test data so that we can hopefully can make this public after the first iteration of the shared task finished, but this is not ready at the moment.)
5.) Yes, the idea is that we run the same submitted Docker image against different hosted LLM endpoints without changing the configuration.
Does this answer your questions?
And a big thank you for taking care about those aspects!
Best regards,
Maik
1 Like
Thanks, this answers our questions. One final clarification: after uploading a Docker submission to TIRA, can participants currently trigger a test execution with the organizer-hosted LLM environment themselves, or are those hosted-model runs only performed by the organizers during the official evaluation?