Jacob Johnson jacobkj314 - Bluesky Statics

Screenshot of the abstract for an academic paper "How Much Consistency Is Your Accuracy Worth?" by Jacob K Johnson and Ana Marasović "Contrast set consistency is a robustness measurement that evaluates the rate at which a model correctly responds to all instances in a bundle of minimally different examples relying on the same knowledge. To draw additional insights, we propose to complement consistency with relative consistency -- the probability that an equally accurate model would surpass the consistency of the proposed model, given a distribution over possible consistencies. Models with 100% relative consistency have reached a consistency peak for their accuracy. We reflect on prior work that reports consistency in contrast sets and observe that relative consistency can alter the assessment of a model's consistency compared to another. We anticipate that our proposed measurement and insights will influence future studies aiming to promote consistent behavior in models."

Some insightful recent works report model consistency across bundles of related instances. But since this naturally increases with accuracy, how should these consistency scores at different accuracies be compared? Our paper for BlackboxNLP@EMNLP2023: arxiv.org/abs/2310.13781

31.10.2023 20:34 — 👍 2 🔁 0 💬 0 📌 1

Posts by Jacob Johnson (@jacobkj314.bsky.social)