Reasoning Trajectory Detection 2026
Synopsis
- Task: Given a triplet (user query, reasoning trajectory, final answer), detect (1) the source of the reasoning/answer (human vs. AI) and (2) the safety of the reasoning and the final answer.
-
Subtasks:
- Subtask 1 (Source Detection): Identify whether the reasoning trajectory and final answer are human-written or AI-generated.
- Subtask 2 (Safety Detection): Decide whether the reasoning trajectory is safe vs. unsafe, and whether the final answer is safe vs. unsafe.
- Important dates:
- February 12, 2026: Training + Validation set release
- March 30, 2026: Testing set release
May 07, 2026May 21, 2026: Prediction submission- May 28, 2026: Participant notebook submission [template] [submission – select "Stylometry and Digital Text Forensics (PAN)" ]
- Data: Accessible at our GitHub repository.
- Evaluation: File-based submissions on a separate evaluation platform (not executed inside the TIRA sandbox).
- Participation: 16 participating teams; nine evaluated systems for Subtask 1 and seven for Subtask 2. Seven teams submitted notebook papers, of which six were accepted.
- Best results: 0.85 macro-F1 for Source Detection; 0.75 trajectory-level F1 and 0.31 step-level soft-F1 for Safety Detection.
Task Overview
The year 2025 saw major advances in the reasoning capabilities of large language models, where models produce explicit reasoning trajectories before a final answer. However, intermediate reasoning steps can still be spurious, non-logical, or unsafe, and in some cases models may reach safe conclusions through deceptive or misaligned reasoning paths. To deepen our understanding of LLM-generated reasoning and to support improvements in reasoning and safety, this task focuses on detecting the source and the safety of reasoning trajectories.
Subtask 1: Source Detection
Given a triplet (user query, reasoning trajectory, final answer), identify whether the reasoning trajectory and final answer are generated by an AI system or written by a human. Queries in the testing set may involve math, coding, and real-life financial reasoning tasks.
Subtask 2: Safety Detection
Given a triplet (user query, reasoning trajectory, final answer), classify (1) whether the reasoning trajectory (i.e. each step in the reasoning trace) is safe vs. unsafe and (2) whether the final answer is safe vs. unsafe. The user queries come from three categories:
- (a) risky queries requesting harmful content,
- (b) jailbreak attacks with risks obscured by various strategies,
- (c) benign queries containing risky tokens.
Data
Dataset release details (access, licensing, and download links) is announced at our GitHub repository.
| Subtask | Train | Validation | Test |
|---|---|---|---|
| Source Detection | 87,556 | 497 | 2,617 |
| Safety Detection | 6,079 | 2,200 | 2,832 |
Source Detection Data
Each instance contains a problem statement, its solution, a coarse source label, and a detailed generator label. Human training solutions are drawn from mathematical reasoning corpora, including the AOPS subset of NuminaMath-CoT and Nemotron. The released LLM solutions were generated using GPT-5 Nano, Gemini-3 Flash, K2-Think V2, and DeepSeek R1.
The test set contains 873 human and 1,744 LLM solutions across 2,083 unique problems. It includes 18 LLM generators, 14 of which are absent from training and validation. These additional generators account for 1,034 test examples. The broader domains, multilingual problems, and variable solution lengths test whether source cues transfer beyond the training distribution.
Safety Detection Data
The data are based on ReasoningShield. Each instance includes the user query, a segmented reasoning trace, a global safety label, and an aligned vector of binary step labels. Test trajectories were manually annotated at both levels.
| Label | Train | Validation | Test |
|---|---|---|---|
| Safe | 5,150 | 1,238 | 1,281 |
| Potentially unsafe | 41 | 334 | 491 |
| Unsafe | 888 | 628 | 1,060 |
The test set contains 10,312 annotated steps, with a median of four steps per trajectory; 30.9% of steps are unsafe. Only 40.3% of test queries are English or mostly English. The trajectory label is not determined simply by whether any unsafe step appears, so the two prediction levels must be evaluated separately.
Submission
Submissions were prediction files in CSV format. Participants ran their systems on their own hardware or services, including local or cloud GPUs, open-weight models, external APIs, or rule-based systems. Evaluation did not require execution inside the TIRA sandbox.
The official repository specifies the filename submission.csv and provides a starter kit for the exact output schema and submission procedure. The competition pages are:
Evaluation
Subtask 1: Source Detection
Systems are ranked by macro-averaged F1 over the human and LLM classes. Accuracy is a secondary metric. Giving both classes equal weight limits the effect of their unequal frequencies.
Macro-F1 = (F1(human) + F1(LLM)) / 2.
Subtask 2: Safety Detection
The primary ranking metric is trajectory-level macro-F1 after mapping the three released labels to two evaluation classes:
| Released label | Evaluation class |
|---|---|
| Safe | Safe |
| Potentially unsafe | Unsafe |
| Unsafe | Unsafe |
Trajectory-F1 = (F1(safe) + F1(unsafe)) / 2.
The secondary metric, step-level soft-F1, measures unsafe-step localization. For each trajectory, an F1 score is computed between its predicted and gold step labels; the final score is the mean over trajectories. This is an average of per-trajectory scores, rather than one F1 computed after pooling all steps.
Step F1 for one trajectory = 2 × TPsoft / (2 × TPsoft + FPsoft + FNsoft).
For each aligned step, let p be the predicted unsafe score and y the gold binary label. The soft counts sum p × y for TPsoft, p × (1 − y) for FPsoft, and (1 − p) × y for FNsoft. The metric supports soft predictions; many systems submitted hard 0/1 step labels. Use the official evaluation code for implementation details and edge cases.
Baselines
The official baselines use supervised fine-tuning of RoBERTa. For Source Detection, RoBERTa is trained as a sequence classifier. For Safety Detection, each segmented trace is decomposed into step examples, the classifier predicts binary step labels, and the predictions are aggregated back into the trace format. Training and inference code are available in the baseline directory.
- Source Detection: 0.55 macro-F1 and 0.57 accuracy.
- Safety Detection: 0.57 trajectory-F1 and 0.24 step-level soft-F1.
Final Leaderboards
Scores below reproduce Tables 7 and 8 of the overview report, rounded to two decimal places. Baselines are shown for reference and are not ranked participant systems. A missing notebook does not mean that no prediction submission was evaluated.
Subtask 1: Source Detection
| Rank | Team | Macro-F1 | Accuracy | Notebook | Approach |
|---|---|---|---|---|---|
| 1 | Writerslogic | 0.85 | 0.85 | Accepted | Claude Opus and Sonnet agreement with LightGBM fallback. |
| 2 | srikarkashyap | 0.83 | 0.82 | Not submitted | Not reported. |
| 3 | DUAN | 0.80 | 0.83 | Accepted | Qwen3.5-2B full fine-tuning with a binary head. |
| 4 | TUKE-AI | 0.75 | 0.74 | Accepted | RoBERTa-base with chunked problem-solution processing. |
| 5 | 23020073 | 0.72 | 0.72 | Not submitted | Not reported. |
| 6 | Asdkkllk | 0.71 | 0.75 | Accepted | ModernBERT-base with solution-only input and a calibrated threshold. |
| 7 | threshdsnmj | 0.65 | 0.67 | Accepted | RoBERTa-base with a short solution-only input. |
| 8 | sajayrrr | 0.43 | 0.43 | Not submitted | Not reported. |
| 9 | krrag | 0.30 | 0.28 | Not submitted | Not reported. |
| Unranked | Baseline | 0.55 | 0.57 | Not applicable | RoBERTa sequence classifier. |
Subtask 2: Safety Detection
| Rank | Team | Trajectory-F1 | Step soft-F1 | Notebook | Approach |
|---|---|---|---|---|---|
| 1 | Bit-by-Bit | 0.75 | 0.31 | Accepted | Seven-criterion Granite Guardian ensemble and a Qwen3Guard step expert. |
| 2 | srikarkashyap | 0.70 | 0.17 | Not submitted | Not reported. |
| 3 | Writerslogic | 0.66 | 0.27 | Accepted | Query harmfulness, multilingual refusal detection, and multi-signal union. |
| 4 | DUAN | 0.65 | 0.16 | Accepted | Qwen3.5-2B LoRA multi-task classifier. |
| 5 | threshdsnmj | 0.56 | 0.09 | Accepted | XLM-RoBERTa trace classifier with weak step supervision. |
| 6 | TUKE-AI | 0.44 | 0.12 | Accepted | XLM-RoBERTa step classifier and BiLSTM-attention trace classifier. |
| 7 | krrag | 0.43 | 0.07 | Not submitted | Not reported. |
| Unranked | Baseline | 0.57 | 0.24 | Not applicable | RoBERTa with step-level decomposition. |
Key Findings
- Generalization matters. Strong validation scores often dropped on the hidden test sets, which introduced changes in domains, generators, languages, trace lengths, and class balance.
- Evidence selection depends on the task. Compact solution-only views were competitive for source detection. Safety detection benefited from the relationship between user intent, reasoning, and refusal or compliance.
- Global safety and local safety remain distinct challenges. Four of seven safety systems exceeded the trajectory-F1 baseline, but only two exceeded the step-level soft-F1 baseline. The best step score was 0.31.
- Decomposition and calibration were recurring strengths. Successful systems separated decisions such as query harmfulness and refusal, or trajectory judgment and step localization. Thresholds tuned on validation did not always transfer to test.
Overview Report
Minh Ngoc Ta, Kaiyang Wan, Yan Lin, Yuxia Wang, and Preslav Nakov. 2026. Overview of the Reasoning Trajectory Detection Task at PAN 2026. CLEF 2026 Working Notes, September 21-24, 2026, Jena, Germany.
Related Work
- Changyi Li et al. 2025. FAID: Fine-Grained AI-Generated Text Detection Using Multi-Task Auxiliary and Multi-Level Contrastive Learning. arXiv [Cs.CL]
- Raj Vardhan Tomar et al. 2025. UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, pages 1233-1247, Mumbai, India.
- Yichi Zhang et al. 2025. STAIR: Improving Safety Alignment with Introspective Reasoning. In Forty-second International Conference on Machine Learning.
- Changyi Li et al. 2025. ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models. arXiv [Cs.CL]



