Are Large Language Models Robust in Understanding Code Against Semantics-Preserving Mutations?

📰 ArXiv cs.AI

arXiv:2505.10443v3 Announce Type: replace-cross Abstract: With the widespread adoption of vibe coding, understanding the reasoning and robustness of Large Language Models (LLMs) is critical for their reliable use in programming tasks. While recent studies assess LLMs' ability to predict program outputs, most focus on accuracy alone, without evaluating the underlying reasoning. Moreover, it has been observed on mathematical reasoning tasks that LLMs can arrive at correct answers through flawed lo

Published 9 May 2026

Full Article

Title: Are Large Language Models Robust in Understanding Code Against Semantics-Preserving Mutations?

Abstract:
arXiv:2505.10443v3 Announce Type: replace-cross Abstract: With the widespread adoption of vibe coding, understanding the reasoning and robustness of Large Language Models (LLMs) is critical for their reliable use in programming tasks. While recent studies assess LLMs' ability to predict program outputs, most focus on accuracy alone, without evaluating the underlying reasoning. Moreover, it has been observed on mathematical reasoning tasks that LLMs can arrive at correct answers through flawed lo
Read full paper → ← Back to Reads