Synopsis

  • Task: Given a triplet (user query, reasoning trajectory, final answer), detect (1) the source of the reasoning/answer (human vs. AI) and (2) the safety of the reasoning and the final answer.
  • Subtasks:
    • Subtask 1 (Source Detection): Identify whether the reasoning trajectory and final answer are human-written or AI-generated.
    • Subtask 2 (Safety Detection): Decide whether the reasoning trajectory is safe vs. unsafe, and whether the final answer is safe vs. unsafe.
  • Important dates:
    • February 12, 2026: Training + Validation set release
    • March 30, 2026: Testing set release
    • May 07, 2026 May 21, 2026: Prediction submission
    • May 28, 2026: Participant notebook submission [template] [submission  – select "Stylometry and Digital Text Forensics (PAN)" ]
  • Data: Accessible at our GitHub repository.
  • Evaluation: File-based submissions on a separate evaluation platform (not executed inside the TIRA sandbox).
  • Participation: 16 participating teams; nine evaluated systems for Subtask 1 and seven for Subtask 2. Seven teams submitted notebook papers, of which six were accepted.
  • Best results: 0.85 macro-F1 for Source Detection; 0.75 trajectory-level F1 and 0.31 step-level soft-F1 for Safety Detection.

Task Overview

The year 2025 saw major advances in the reasoning capabilities of large language models, where models produce explicit reasoning trajectories before a final answer. However, intermediate reasoning steps can still be spurious, non-logical, or unsafe, and in some cases models may reach safe conclusions through deceptive or misaligned reasoning paths. To deepen our understanding of LLM-generated reasoning and to support improvements in reasoning and safety, this task focuses on detecting the source and the safety of reasoning trajectories.

Subtask 1: Source Detection

Given a triplet (user query, reasoning trajectory, final answer), identify whether the reasoning trajectory and final answer are generated by an AI system or written by a human. Queries in the testing set may involve math, coding, and real-life financial reasoning tasks.

Subtask 2: Safety Detection

Given a triplet (user query, reasoning trajectory, final answer), classify (1) whether the reasoning trajectory (i.e. each step in the reasoning trace) is safe vs. unsafe and (2) whether the final answer is safe vs. unsafe. The user queries come from three categories:

  • (a) risky queries requesting harmful content,
  • (b) jailbreak attacks with risks obscured by various strategies,
  • (c) benign queries containing risky tokens.

Data

Dataset release details (access, licensing, and download links) is announced at our GitHub repository.

Dataset sizes from Table 1 of the overview report.
SubtaskTrainValidationTest
Source Detection87,5564972,617
Safety Detection6,0792,2002,832

Source Detection Data

Each instance contains a problem statement, its solution, a coarse source label, and a detailed generator label. Human training solutions are drawn from mathematical reasoning corpora, including the AOPS subset of NuminaMath-CoT and Nemotron. The released LLM solutions were generated using GPT-5 Nano, Gemini-3 Flash, K2-Think V2, and DeepSeek R1.

The test set contains 873 human and 1,744 LLM solutions across 2,083 unique problems. It includes 18 LLM generators, 14 of which are absent from training and validation. These additional generators account for 1,034 test examples. The broader domains, multilingual problems, and variable solution lengths test whether source cues transfer beyond the training distribution.

Safety Detection Data

The data are based on ReasoningShield. Each instance includes the user query, a segmented reasoning trace, a global safety label, and an aligned vector of binary step labels. Test trajectories were manually annotated at both levels.

Trajectory-level safety labels from Table 4 of the overview report.
LabelTrainValidationTest
Safe5,1501,2381,281
Potentially unsafe41334491
Unsafe8886281,060

The test set contains 10,312 annotated steps, with a median of four steps per trajectory; 30.9% of steps are unsafe. Only 40.3% of test queries are English or mostly English. The trajectory label is not determined simply by whether any unsafe step appears, so the two prediction levels must be evaluated separately.

Submission

Submissions were prediction files in CSV format. Participants ran their systems on their own hardware or services, including local or cloud GPUs, open-weight models, external APIs, or rule-based systems. Evaluation did not require execution inside the TIRA sandbox.

The official repository specifies the filename submission.csv and provides a starter kit for the exact output schema and submission procedure. The competition pages are:

Evaluation

Subtask 1: Source Detection

Systems are ranked by macro-averaged F1 over the human and LLM classes. Accuracy is a secondary metric. Giving both classes equal weight limits the effect of their unequal frequencies.

Macro-F1 = (F1(human) + F1(LLM)) / 2.

Subtask 2: Safety Detection

The primary ranking metric is trajectory-level macro-F1 after mapping the three released labels to two evaluation classes:

Label mapping for the primary Safety Detection score.
Released labelEvaluation class
SafeSafe
Potentially unsafeUnsafe
UnsafeUnsafe

Trajectory-F1 = (F1(safe) + F1(unsafe)) / 2.

The secondary metric, step-level soft-F1, measures unsafe-step localization. For each trajectory, an F1 score is computed between its predicted and gold step labels; the final score is the mean over trajectories. This is an average of per-trajectory scores, rather than one F1 computed after pooling all steps.

Step F1 for one trajectory = 2 × TPsoft / (2 × TPsoft + FPsoft + FNsoft).

For each aligned step, let p be the predicted unsafe score and y the gold binary label. The soft counts sum p × y for TPsoft, p × (1 − y) for FPsoft, and (1 − p) × y for FNsoft. The metric supports soft predictions; many systems submitted hard 0/1 step labels. Use the official evaluation code for implementation details and edge cases.

Baselines

The official baselines use supervised fine-tuning of RoBERTa. For Source Detection, RoBERTa is trained as a sequence classifier. For Safety Detection, each segmented trace is decomposed into step examples, the classifier predicts binary step labels, and the predictions are aggregated back into the trace format. Training and inference code are available in the baseline directory.

  • Source Detection: 0.55 macro-F1 and 0.57 accuracy.
  • Safety Detection: 0.57 trajectory-F1 and 0.24 step-level soft-F1.

Final Leaderboards

Scores below reproduce Tables 7 and 8 of the overview report, rounded to two decimal places. Baselines are shown for reference and are not ranked participant systems. A missing notebook does not mean that no prediction submission was evaluated.

Subtask 1: Source Detection

Nine evaluated systems, ranked by macro-F1. Higher scores are better.
RankTeamMacro-F1AccuracyNotebookApproach
1Writerslogic0.850.85AcceptedClaude Opus and Sonnet agreement with LightGBM fallback.
2srikarkashyap0.830.82Not submittedNot reported.
3DUAN0.800.83AcceptedQwen3.5-2B full fine-tuning with a binary head.
4TUKE-AI0.750.74AcceptedRoBERTa-base with chunked problem-solution processing.
5230200730.720.72Not submittedNot reported.
6Asdkkllk0.710.75AcceptedModernBERT-base with solution-only input and a calibrated threshold.
7threshdsnmj0.650.67AcceptedRoBERTa-base with a short solution-only input.
8sajayrrr0.430.43Not submittedNot reported.
9krrag0.300.28Not submittedNot reported.
UnrankedBaseline0.550.57Not applicableRoBERTa sequence classifier.

Subtask 2: Safety Detection

Seven evaluated systems, ranked by trajectory-F1. Soft-F1 is the secondary metric; higher scores are better.
RankTeamTrajectory-F1Step soft-F1NotebookApproach
1Bit-by-Bit0.750.31AcceptedSeven-criterion Granite Guardian ensemble and a Qwen3Guard step expert.
2srikarkashyap0.700.17Not submittedNot reported.
3Writerslogic0.660.27AcceptedQuery harmfulness, multilingual refusal detection, and multi-signal union.
4DUAN0.650.16AcceptedQwen3.5-2B LoRA multi-task classifier.
5threshdsnmj0.560.09AcceptedXLM-RoBERTa trace classifier with weak step supervision.
6TUKE-AI0.440.12AcceptedXLM-RoBERTa step classifier and BiLSTM-attention trace classifier.
7krrag0.430.07Not submittedNot reported.
UnrankedBaseline0.570.24Not applicableRoBERTa with step-level decomposition.

Key Findings

  • Generalization matters. Strong validation scores often dropped on the hidden test sets, which introduced changes in domains, generators, languages, trace lengths, and class balance.
  • Evidence selection depends on the task. Compact solution-only views were competitive for source detection. Safety detection benefited from the relationship between user intent, reasoning, and refusal or compliance.
  • Global safety and local safety remain distinct challenges. Four of seven safety systems exceeded the trajectory-F1 baseline, but only two exceeded the step-level soft-F1 baseline. The best step score was 0.31.
  • Decomposition and calibration were recurring strengths. Successful systems separated decisions such as query harmfulness and refusal, or trajectory judgment and step localization. Thresholds tuned on validation did not always transfer to test.

Overview Report

Minh Ngoc Ta, Kaiyang Wan, Yan Lin, Yuxia Wang, and Preslav Nakov. 2026. Overview of the Reasoning Trajectory Detection Task at PAN 2026. CLEF 2026 Working Notes, September 21-24, 2026, Jena, Germany.

Task Committee