Skip to content
Research

Bridging the GAP between Computer Vision and Generative AI

We work on problems we find genuinely interesting in Computer Vision and Generative AI for both generation and perception and their intersection.

Research Areas

Diffusion Models & Generative Synthesis

Diffusion Models & Generative Synthesis

Controllable image and video generation with diffusion models — from fine-grained, part-level editing to consistent character animation and 3D-aware scene layout.

Multi-Modal LLMs & Agentic Systems

Multi-Modal LLMs & Agentic Systems

Vision-language models and agentic pipelines that reason over images, diagrams and video — not just generate text — and act on that reasoning.

Retrieval-Augmented Generation

Retrieval-Augmented Generation

Grounding generative models in real, verifiable knowledge — with evaluation harnesses that measure reliability instead of assuming it.

3D Vision & Scene Understanding

3D Vision & Scene Understanding

Shape correspondence, keypoint reasoning, and language-guided object placement in real 3D scenes — connecting geometric understanding with natural-language interaction.

Uncertainty-Aware Perception

Uncertainty-Aware Perception

Confidence-propagating networks for sparse and noisy data — depth completion, optical flow, and regression tasks where knowing what the model doesn't know matters.

Detection, Tracking & Video Segmentation

Detection, Tracking & Video Segmentation

Robust object detection, tracking and segmentation in real video — including distractor-aware tracking and zero-shot segmentation built on pre-trained diffusion models.

Publications

2025

MV
NeurIPS 2025Spotlight

Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation

Abdelrahman Eldesokey, Aleksandar Cvejic, Bernard Ghanem, Peter Wonka

Introduces a method for detecting visual inconsistencies in subject-driven image generation by leveraging visual correspondence, improving reliabil…

ER
ICCV 2025

EditCLIP: Representation Learning for Image Editing

Qian Wang, Aleksandar Cvejic, Abdelrahman Eldesokey, Peter Wonka

A representation-learning approach tailored to image editing, learning embeddings that capture the semantics of an edit itself rather than just ima…

PL
ICCV 2025

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

Ahmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi, Abdelrahman Eldesokey, Peter Wonka, Gabriel Brostow, Sara Vicente, Guillermo Garcia-Hernando

Enables placing objects into real 3D scenes using natural-language instructions, bridging language understanding with geometric scene reasoning.

ZP
ICCV 2025

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

Bingchen Gong, Diego Gomez, Abdullah Hamdi, Abdelrahman Eldesokey, Ahmed Abdelreheem, Peter Wonka, Maks Ovsjanikov

Shows that large language models can be used for zero-shot, point-level reasoning to detect semantic 3D keypoints without task-specific training da…

PF
SIGGRAPH 2025

PartEdit: Fine-Grained Image Editing using Pre-Trained Diffusion Models

Aleksandar Cvejic*, Abdelrahman Eldesokey*, Peter Wonka

A method for precise, part-level image editing built on pre-trained diffusion models, allowing targeted edits to specific object parts without dist…

VZ
CVPR 2025

VidSeg: Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models

Qian Wang, Abdelrahman Eldesokey, Mohit Mendiratta, Fangneng Zhan, Adam Kortylewski, Christian Theobalt, Peter Wonka

Repurposes pre-trained diffusion models for zero-shot semantic segmentation of video, removing the need for task-specific labeled training data.

BI
ICLR 2025

Build-A-Scene: Interactive 3D Layout Control for Diffusion-Based Image Generation

Abdelrahman Eldesokey, Peter Wonka

Gives users interactive, explicit control over 3D object layout when generating images with diffusion models, closing the gap between free-form pro…