Foundations
Computer Vision
Object detection, segmentation, YOLO, CLIP, and vision-language models
Skills in this topic
3 skills — Sign in to track your progress
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
1d ago
What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation
arXiv:2608.21761v1 Announce Type: new Abstract: Large collections of street-view imagery provide rich visual information about urban environments, but extractin
Weights & Biases
👁️ Computer Vision
⚡ AI Lesson
2w ago
When axis-aligned boxes break: Lessons from a CVPR-published traffic AI
A case study in treating representation as a design decision, not a given — with the code and the logging to prove it.
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
arXiv:2608.05218v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
arXiv:2608.06150v1 Announce Type: new Abstract: Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories.
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
arXiv:2608.05745v1 Announce Type: cross Abstract: Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
arXiv:2608.05771v1 Announce Type: cross Abstract: Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation,
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
arXiv:2608.06075v1 Announce Type: cross Abstract: Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Depth-Guided Video Object Counting in Crowded Scenes
arXiv:2608.06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all inst
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
arXiv:2608.06240v1 Announce Type: cross Abstract: Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired
Reddit r/MachineLearning
👁️ Computer Vision
⚡ AI Lesson
2w ago
[R], Need some best model suggestions for Face Detection,Face Recognition,Body Detection and Body identification. [R]
need those for analysing movies. example let's say I have to find the screentime of the actor over the whole runtime of the movie and i need to do it for the
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers
arXiv:2608.04035v1 Announce Type: cross Abstract: The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns re
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
arXiv:2608.04124v1 Announce Type: cross Abstract: Video question answering requires models to ground language queries in visual evidence and, when necessary, re
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation
arXiv:2608.04154v1 Announce Type: cross Abstract: Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because ter
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
arXiv:2608.04589v1 Announce Type: cross Abstract: EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimod
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion
arXiv:2608.04655v1 Announce Type: cross Abstract: Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable g
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Towards a satellite image manipulation and deepfake localization benchmark dataset
arXiv:2608.04840v1 Announce Type: cross Abstract: Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection
arXiv:2608.05069v1 Announce Type: cross Abstract: Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
arXiv:2608.05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computat
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
ZoomV: Temporal Zoom-in for Efficient Long Video Understanding
arXiv:2504.01407v3 Announce Type: replace-cross Abstract: Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
arXiv:2511.19418v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual unde
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
2w ago
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
arXiv:2603.06140v2 Announce Type: replace-cross Abstract: Video object insertion is fundamental to video editing, yet existing diffusion methods often produce v
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Multi-Camera Trajectory Forecasting with Trajectory Tensors
arXiv:2108.04694v2 Announce Type: cross Abstract: We introduce the problem of multi-camera trajectory forecasting (MCTF), which involves predicting the trajecto
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
arXiv:2403.09281v3 Announce Type: cross Abstract: We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation. While the CLIP mo
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
arXiv:2608.03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
arXiv:2608.03127v1 Announce Type: cross Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand gene
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region
arXiv:2608.03407v1 Announce Type: cross Abstract: Road network segmentation from satellite imagery remains challenging due to large geographic variation in road
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Multi-Task Multi-Frame Visual Piano Transcription
arXiv:2608.03419v1 Announce Type: cross Abstract: Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
arXiv:2608.03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained f
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
arXiv:2608.03918v1 Announce Type: cross Abstract: Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of fr
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Compound and Parallel Modes of Tropical Convolutional Neural Networks
arXiv:2504.06881v2 Announce Type: replace-cross Abstract: Convolutional neural networks (CNNs) are foundational to many state-of-the-art computer vision systems
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
ORCA: ORgan-Centroid Aggregation for Training-Free 3D CT Visual Token Compression
arXiv:2608.00345v1 Announce Type: cross Abstract: A 3D CT scan entering a vision-language model produces a long sequence of visual tokens, often thousands to te
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Boosting Generalizable Depth Estimation in Endoscopy by Mixture of Lightweight Experts and Intrinsic Image Alignment
arXiv:2608.00415v1 Announce Type: cross Abstract: Depth estimation is a significant task for 3D perception in endoscopic surgeries. However, illumination interf
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
RadYOLO: Computationally Efficient 3D Object Detection and Segmentation in CT and MRI
arXiv:2608.00508v1 Announce Type: cross Abstract: Object detection and segmentation in three-dimensional medical images is a very active area of research. Howev
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations
arXiv:2608.00530v1 Announce Type: cross Abstract: Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Coverage-Driven Adaptive Keyframe Selection for Video Understanding
arXiv:2608.00714v1 Announce Type: cross Abstract: Recent advances in large vision-language models (LVLMs) have enabled long-video understanding and analysis. Ho
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors
arXiv:2608.00946v1 Announce Type: cross Abstract: Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis o
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction
arXiv:2608.01178v1 Announce Type: cross Abstract: We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS)
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment
arXiv:2608.01301v1 Announce Type: cross Abstract: Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked b
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
VGER: Voxel-Guided Global Event Ranking for Event Cloud Attribution
arXiv:2608.01470v1 Announce Type: cross Abstract: Event cameras produce sparse and asynchronous event streams that provide rich spatio-temporal information for
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Enhancing Visual Perception in Foggy Conditions via Multiclass Fog Density Modeling
arXiv:2608.01572v1 Announce Type: cross Abstract: Autonomous driving (AD) systems have advanced rapidly over the past decade; however, robust perception under a
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
arXiv:2608.01643v1 Announce Type: cross Abstract: Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streami
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Entity-Aware Sequence Transduction for Player-Centric Ball Action Spotting
arXiv:2608.01696v1 Announce Type: cross Abstract: Player-centric ball action spotting requires temporally precise event detection together with actor attributio
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
When Extreme Darkness Meets Motion Blur: MeanFlow for Unified RAW Restoration
arXiv:2608.01720v1 Announce Type: cross Abstract: Extremely low-light RAW enhancement aims to recover severely attenuated sensor signals, yet existing methods o
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit
arXiv:2608.01753v1 Announce Type: cross Abstract: Addressing urban blight has seen increased focus in the past 15 years. Assessing urban blight is essential for
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis
arXiv:2608.01958v1 Announce Type: cross Abstract: 4D Gaussian Splatting (4DGS) excels in dynamic 3D reconstruction and real-time novel view synthesis via effici
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
TBSG-Net: Temporal Bipartite Scene Graph Network for Fine-Grained Video Moment Retrieval
arXiv:2608.02056v1 Announce Type: cross Abstract: Recent advances in proposal-free Video Moment Retrieval (VMR) have highlighted the effectiveness of Static Sce
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization
arXiv:2608.02099v1 Announce Type: cross Abstract: 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existi
ArXiv cs.AI
👁️ Computer Vision
📄 Paper
⚡ AI Lesson
3w ago
UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction
arXiv:2608.02145v1 Announce Type: cross Abstract: In this paper, we propose UniqueSplat, a view-conditioned feed-forward 3D Gaussian Splatting model to reconstr
DeepCamp AI