HCML Reading Group Session on 24.09.2024

50 views

Skip to first unread message

Aditya Gulati

unread,

Sep 12, 2024, 7:53:17 AM9/12/24

to ELLIS-Human-Centric ML

Hello!

Hope you had a relaxing and overall wonderful summer :) Here are the details of the next session of our reading group:

Title: Describing Differences in Image Sets with Natural Language

Authors: Lisa Dunlap, Yuhui Zhang, Xiaohan Wang, Ruiqi Zhong, Trevor Darrell, Jacob Steinhardt, Joseph E. Gonzalez, Serena Yeung-Levy

Abstract: How do two sets of images differ? Discerning set-level differences is crucial for understanding model behaviors and analyzing datasets yet manually sifting through thousands of images is impractical. To aid in this discovery process we explore the task of automatically describing the differences between two sets of images which we term Set Difference Captioning. This task takes in image sets \mathcal D _A and \mathcal D _B and outputs a description that is more often true on \mathcal D _A than \mathcal D _B. We outline a two-stage approach that first proposes candidate difference descriptions from image sets and then re-ranks the candidates by checking how well they can differentiate the two sets. We introduce VisDiff which first captions the images and prompts a language model to propose candidate descriptions then re-ranks these descriptions using CLIP. To evaluate VisDiff we collect VisDiffBench a dataset with 187 paired image sets with ground truth difference descriptions. We apply VisDiff to various domains such as comparing datasets (e.g. ImageNet vs. ImageNetV2) comparing classification models (e.g. zero-shot CLIP vs. supervised ResNet) characterizing differences between generative models (e.g. StableDiffusionV1 and V2) and discovering what makes images memorable. Using VisDiff we are able to find interesting and previously unknown differences in datasets and models demonstrating its utility in revealing nuanced insights.

Paper link: https://openaccess.thecvf.com/content/CVPR2024/html/Dunlap_Describing_Differences_in_Image_Sets_with_Natural_Language_CVPR_2024_paper.html

Presenter: Piera Riccio

When: 24th of September at 3pm CEST
Where: usual meeting link! :D

Looking forward to seeing you there!

Best,
Aditya

Piera Riccio

unread,

Sep 24, 2024, 9:01:04 AM9/24/24

to ELLIS-Human-Centric ML

Hello everyone! The session is about to start :)

Reply all

Reply to author

Forward

0 new messages