Core AI

Large Language Models

Deep dives into GPT, Claude, Gemini, Llama and the transformers powering modern AI

74,739
lessons
Skills in this topic
View full skill map →
LLM Foundations
beginner
Explain how transformers generate text
Prompt Craft
beginner
Write zero-shot and few-shot prompts
LLM Engineering
intermediate
Call LLM APIs with function/tool use
Fine-tuning LLMs
advanced
Prepare fine-tuning datasets
Multimodal LLMs
advanced
Use GPT-4V / Claude Vision for image understanding
All Reads (51,330) Articles (21740)Blog Posts (9493)Tutorials (4485)Research Papers (14164)News (1448)
Reddit r/MachineLearning 🧠 Large Language Models ⚡ AI Lesson 1d ago
A dataset with 52 Text to image model evaluation [P]
I created a simple text to image benchmark. I curated 192 prompts that are difficult for T2I models in various ways: text rendering, spatial reasoning, human re
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
LLM Agents Perform Controlled Experiments Using Simulation Models
arXiv:2608.23622v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scien
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
TRACE: Transition-Aware Residual Control for Multi-Objective Materials Discovery
arXiv:2608.23631v1 Announce Type: new Abstract: Multi-objective materials discovery with LLM agents is often limited not only by how many candidates can be prop
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes
arXiv:2608.23640v1 Announce Type: new Abstract: When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happene
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
How much of a measured AI preference is the model, and how much is the instrument?
arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit prefer
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value
arXiv:2608.23644v1 Announce Type: new Abstract: Large language models (LLMs) are becoming routine instruments of scientific research, assisting with literature
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models
arXiv:2608.23646v1 Announce Type: new Abstract: Molecular embedding models can serve as foundational infrastructure for computational chemistry and drug discove
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Gated Activation Steering for Reducing Sycophancy & Hallucination in Medical Question Answering
arXiv:2608.23666v1 Announce Type: new Abstract: Sycophancy and hallucination are persistent failure modes of Large Language Models (LLMs) across domains. Howeve
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Do LLMs Understand Limit Order Book Dynamics?
arXiv:2608.23706v1 Announce Type: new Abstract: A large language model (LLM) trained on synthetic limit order book (LOB) data achieves near perfect scores in ge
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Serving Masked Diffusion LLMs: Characterization and Design Principles from Real Hardware
arXiv:2608.23807v1 Announce Type: new Abstract: Masked diffusion language models (dLLMs) can in principle generate text faster than autoregressive (AR) models,
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Generating Biomedical Fact-Checking Reports with RL-Enhanced Agentic Search
arXiv:2608.23811v1 Announce Type: new Abstract: Automated fact-checking is essential for ensuring the reliability of public health information, yet the biomedic
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention
arXiv:2608.23834v1 Announce Type: new Abstract: The key-value (KV) cache is a primary capacity and bandwidth bottleneck in long-context LLM serving. We present
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models
arXiv:2608.23837v1 Announce Type: new Abstract: Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with users in
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications
arXiv:2608.23870v1 Announce Type: new Abstract: When it comes to safety policies for generative AI, one size does not fit all. Each organization and use case ne
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors
arXiv:2608.23873v1 Announce Type: new Abstract: Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
AI Finds A Way
arXiv:2608.23875v1 Announce Type: new Abstract: Artificial Intelligence (AI) algorithms frequently learn creative and unexpected solutions, surprising even expe
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 1d ago
BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification
arXiv:2608.23898v1 Announce Type: new Abstract: We introduce BenchBench-Protocol, a benchmark for large language models of 149 protocol-modification tasks recov
I spent a day at a robot “carnival” in Shanghai. Here’s what I saw.
MIT Technology Review 🧠 Large Language Models ⚡ AI Lesson 2d ago
I spent a day at a robot “carnival” in Shanghai. Here’s what I saw.
Humanoid robots are having a moment in China. The popular machines are part of the country’s strategy to bring artificial intelligence into daily life. Embeddin
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
AI Learning and Conceptual Transfer in the Game of Hidden Rules
arXiv:2608.21372v1 Announce Type: new Abstract: This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG
arXiv:2608.21375v1 Announce Type: new Abstract: Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, inter
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items
arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the opti
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference
arXiv:2608.21393v1 Announce Type: new Abstract: Running large language models inside enterprise environments has always bumped up against a practical wall: the
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering
arXiv:2608.21408v1 Announce Type: new Abstract: Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown expon
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Composable Trust Infrastructure for Manufacturing Knowledge Graphs: Cross-System Provenance, Temporal Reasoning, and Decision Traceability
arXiv:2608.21418v1 Announce Type: new Abstract: Manufacturing knowledge graphs that integrate data from heterogeneous industrial systems face a trust deficit: c
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
arXiv:2608.21430v1 Announce Type: new Abstract: Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Software Frameworks for Explainable AI in Time Series Classification: A Systematic Review
arXiv:2608.21449v1 Announce Type: new Abstract: Time series arise in a wide range of application domains and are analyzed using machine learning in decision-cri
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning
arXiv:2608.21501v1 Announce Type: new Abstract: Credit assignment in large-language-model reinforcement learning (LLM RL) can be separated into three objects: e
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification
arXiv:2608.21570v1 Announce Type: new Abstract: Deploying a safety layer for large language models on commodity hardware is constrained by the guards available
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Data-Driven Dynamic Algorithm Dispatch with Large Language Models
arXiv:2608.21584v1 Announce Type: new Abstract: We introduce a large language model (LLM)-driven approach for generating dynamic algorithmic dispatch heuristics
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Semantic Compression Trees: Multi-Resolution Knowledge Retrieval via Hierarchical Semantic Residuals
arXiv:2608.21610v1 Announce Type: new Abstract: Retrieval-augmented generation relies mostly on flat, fixed-granularity indexes: documents are cut into uniform
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning
arXiv:2608.21614v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps,
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism
arXiv:2608.21702v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal re
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling
arXiv:2608.21712v1 Announce Type: new Abstract: Transformer-based models are widely used for clinical prediction from electronic health records (EHRs), yet thei
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries
arXiv:2608.21792v1 Announce Type: new Abstract: Document classification in regulated industries is constrained by data residency, limited cold-start labels, sca
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
VisAdj: Learning Adjacency Matrices from Node-Link Images
arXiv:2608.21825v1 Announce Type: new Abstract: Learning adjacency matrices from node-link images is a fundamental problem for recovering structured graph infor
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents
arXiv:2608.21830v1 Announce Type: new Abstract: Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong pote
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization
arXiv:2608.21836v1 Announce Type: new Abstract: Large language models have become increasingly capable agents for low-level code and kernel optimization, but is
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance
arXiv:2608.21867v1 Announce Type: new Abstract: LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capabi
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation
arXiv:2608.21925v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeut
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Multimodal Prompt Learning with Irregular EHRs for Robust Monitoring of Critical Care Patients
arXiv:2608.21941v1 Announce Type: new Abstract: Accurate assessment of patients in intensive care units (ICUs) is essential for timely clinical intervention and
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Repo2Skill-Evo: Repository Skills Go Stale in Silence
arXiv:2608.21964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over evolving software repositories, where success depend
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning
arXiv:2608.22048v1 Announce Type: new Abstract: Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons.
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers
arXiv:2608.22063v1 Announce Type: new Abstract: Agents built on Large Language Models (LLMs) increasingly reach enterprise data through the Model Context Protoc
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays
arXiv:2608.22068v1 Announce Type: new Abstract: Geothermal well arrays, which organize multiple geothermal wells into carefully planned geometric configurations
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation
arXiv:2608.22085v1 Announce Type: new Abstract: Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning
arXiv:2608.22128v1 Announce Type: new Abstract: Printability assessment in additive manufacturing is typically conducted at the geometry level before printing t
ArXiv cs.AI 🧠 Large Language Models 📄 Paper ⚡ AI Lesson 2d ago
MegaMem: A Retrieval Solution for Ultra-Large Context Windows
arXiv:2608.22137v1 Announce Type: new Abstract: Modern language models and agents increasingly require persistent memory for complete codebases, long interactio
How Claude Watermarks AI-Generated Text
Ahead of AI 🧠 Large Language Models ⚡ AI Lesson 5d ago
How Claude Watermarks AI-Generated Text
A 48-minute video walkthrough of token sampling, watermark detection, and removal