Abstract
Recent advancements in large language models (LLMs) have substantially improved multi-step reasoning in mathematical problem-solving, particularly through techniques such as Chain-of-Thought (CoT) prompting. Building on this progress, the Wrong-of-Thought (WoT) framework introduces Multi-Perspective Verification (MPV) to validate model reasoning through assertion, process, and result checks, while also introducing the Wrong Information Utilization (WIU) mechanism, which uses prior erroneous reasoning to reduce repeated mistakes. However, WoT still suffers from three critical limitations: (i) MPV remains vulnerable to hallucination, leading to false positive and false negative verification decisions; (ii) WoT is computationally inefficient because logically valid reasoning may be discarded when execution fails due to minor formatting or syntax issues, thereby triggering unnecessary reasoning iterations for these cases; and (iii) the WIU mechanism provides prior failures only as raw equations or code, without clearly identifying the error type or location, which limits their usefulness for subsequent refinement. To address these limitations, we propose the (Adaptive Verifier & Reasoner) AVR framework, which enhances WoT by jointly addressing three coupled failure modes: unreliable verification, ineffective reuse of prior errors, and brittle execution handling. Specifically, AVR introduces hierarchical verification through an outer verifier that aggregates MPV insights and signals to make more reliable final decisions; a diagnosis-guided refinement mechanism that converts prior failed reasoning into natural-language diagnostics describing the error type, location, and cause; and an executor-swapping mechanism that preserves logically valid reasoning by retrying execution with an alternative executor before triggering a new reasoning cycle. Empirical evaluation across three LLMs and six widely used benchmarks demonstrates that AVR consistently outperforms established baselines, achieving higher reasoning accuracy, reducing invalid outputs, and improving computational efficiency and token utilization. Our results show that improving verification reliability, error interpretability, and execution robustness can substantially improve multi-step reasoning in LLMs.
| Original language | English |
|---|---|
| Article number | 132740 |
| Number of pages | 16 |
| Journal | Expert Systems with Applications |
| Volume | 327 |
| DOIs | |
| Publication status | Published - 25 Sept 2026 |
Keywords
- Large language models
- Mathematical reasoning
- Natural-language processing
- Prompt engineering
- Question answering
Fingerprint
Dive into the research topics of 'AVR: An enhanced reasoning framework with adaptive diagnostic verification and natural-language error feedback'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver