Query Regarding the test results

19 views
Skip to first unread message

shivam

unread,
Jul 4, 2026, 1:54:47 AMJul 4
to Touché Workshop Series on Argumentation Systems
Since the results for Subtask-2 have not yet been released on TIRA, what we report in our camera ready paper for subtask-2. 

Maik Fröbe

unread,
Jul 4, 2026, 5:50:26 AMJul 4
to shivam, Touché Workshop Series on Argumentation Systems

Dear Shivam,

Thanks for reaching out!
I am not sure why the results for subtask 2 have not been released, I think this must have been an oversight on our end, sorry for the inconvenience!

I added them to the web page, you can see them here: https://www.tira.io/task-overview/fallacy-detection/

Best regards,

Maik

On 7/4/26 7:54 AM, shivam wrote:
Since the results for Subtask-2 have not yet been released on TIRA, what we report in our camera ready paper for subtask-2.  --
You received this message because you are subscribed to the Google Groups "Touché Workshop Series on Argumentation Systems" group.
To unsubscribe from this group and stop receiving emails from it, send an email to touche-lab+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/touche-lab/609fc899-b46d-439f-9a51-07c8afa28181n%40googlegroups.com.

Konrad Maximilian Heinrich

unread,
Jul 4, 2026, 4:30:12 PMJul 4
to Maik Fröbe, shivam, Touché Workshop Series on Argumentation Systems
Hi Shivam,

You can also find now more detailed results here


Best Regards,

Max

shivam

unread,
Jul 5, 2026, 3:05:20 AMJul 5
to Touché Workshop Series on Argumentation Systems
The result mention on TIRA for Task-2  for  Team="nit-agartala-nlp-team" and approach ="Bagging-LR" is accuracy=0.614, F1-micro=0.614 F1-macro=0.574
 But result that published in Touche website  showing  f1-macro=0.766.
Could you please clarify the difference between these results and confirm which result we should use In paper?

Konrad Maximilian Heinrich

unread,
Jul 5, 2026, 5:32:46 AMJul 5
to shivam, Touché Workshop Series on Argumentation Systems
Dear Shivam,

Good catch — the two numbers come from two different evaluations.

The preliminary TIRA display for Sub-Task 2 scored all 233 arguments against the `resembles_fallacy` field, meaning it also included the 118 non-fallacious arguments. The official task definition, however, evaluates Sub-Task 2 only on the 115 fallacious arguments against `fallacy_type`: “Given a fallacious argument, identify the specific type of fallacy.”

For the fallacious arguments, the two gold fields coincide, so the difference comes entirely from the non-fallacious arguments that were mistakenly included in the automatic Tira evaluation.

Please use the results published on the website in your paper, as these are the official final evaluation results.

Grettings,

Max

Reply all
Reply to author
Forward
0 new messages