Coherence Under Commitment: Probing Generalization and Vacuous Memorization in LLM Logical Reasoning
📰 ArXiv cs.AI
Learn to evaluate LLMs' logical reasoning using Coherence Under Commitment, a dual-query paradigm that measures consistency and decisiveness, to avoid vacuous memorization
Action Steps
- Apply Coherence Under Commitment to your LLM evaluation pipeline to measure consistency and decisiveness
- Use dual-query evaluation to jointly assess entailment and refutation
- Configure your LLM to provide commitment to either entailment or refutation, rather than systematic abstention
- Test your LLM on knowledge-intensive domains to evaluate its logical reasoning capabilities
- Compare the performance of your LLM with and without Coherence Under Commitment to measure the impact on generalization and vacuous memorization
Who Needs to Know This
NLP researchers and engineers working with LLMs can benefit from this approach to improve their models' logical reasoning capabilities, while data scientists and AI engineers can apply this method to evaluate and fine-tune their LLMs
Key Insight
💡 Coherence Under Commitment helps evaluate LLMs' ability to provide meaningful commitments to logical entailments, rather than just satisfying negation consistency
Share This
🤖 Improve LLM logical reasoning with Coherence Under Commitment! 📊 Jointly measure consistency and decisiveness to avoid vacuous memorization #LLMs #NLP
Key Takeaways
Learn to evaluate LLMs' logical reasoning using Coherence Under Commitment, a dual-query paradigm that measures consistency and decisiveness, to avoid vacuous memorization
Full Article
Title: Coherence Under Commitment: Probing Generalization and Vacuous Memorization in LLM Logical Reasoning
Abstract:
arXiv:2606.21083v1 Announce Type: new Abstract: Large language models (LLMs) deployed for logical reasoning in knowledge-intensive domains exhibit a subtle but critical failure: coherence can be vacuously achieved through systematic abstention. A model that withholds commitment to either entailment or refutation satisfies negation consistency while providing no utility. We introduce Coherence Under Commitment (CUC), a dual-query evaluation paradigm that jointly measures consistency and decisiven
Abstract:
arXiv:2606.21083v1 Announce Type: new Abstract: Large language models (LLMs) deployed for logical reasoning in knowledge-intensive domains exhibit a subtle but critical failure: coherence can be vacuously achieved through systematic abstention. A model that withholds commitment to either entailment or refutation satisfies negation consistency while providing no utility. We introduce Coherence Under Commitment (CUC), a dual-query evaluation paradigm that jointly measures consistency and decisiven
DeepCamp AI