Groups
Groups

[CFP: LIGHT Workshop @ NeurIPS Paris 2026] – Lightweight & Trustworthy Foundation Models

17 views
Skip to first unread message

Roberta Calegari

unread,
Jul 23, 2026, 10:47:03 AM (7 days ago) Jul 23
to AIxIA mailing list

Dear colleagues,

We are pleased to announce the Call for Papers for the NeurIPS 2026 Workshop (Paris Venue):

LIGHT – Trustworthy Small Foundation Models
Compact, Trustworthy and Energy-Efficient AI Systems for Real-World Deployment

🌐 Workshop website: https://almaai-disi-unibo.github.io/neurips2026-light-smallModels/

As foundation models continue to grow in size, deploying AI systems in industrial, edge, and regulated environments remains a major challenge. The LIGHT workshop aims to bring together researchers and practitioners working on the next generation of compact, trustworthy, and deployable foundation models, fostering discussion across machine learning, systems, trustworthy AI, neuro-symbolic AI, and industrial deployment.

We welcome submissions on topics including (but not limited to):

  • Distillation of foundation-model capabilities into compact models
  • Small language models and compact multimodal systems
  • Trustworthiness-by-design through model compression
  • Neuro-symbolic extensions for compact foundation models
  • Explainability, reasoning and verification in small models
  • Robustness and safety of compressed AI systems
  • Runtime guarantees for deployable AI
  • Efficient adaptation beyond fine-tuning
  • Edge and resource-constrained deployment
  • Sustainable AI through compact architectures
  • Benchmarks and evaluation methodologies for small trustworthy models
  • Industrial deployment of compact and trustworthy AI

We encourage submissions of extended abstracts, position papers, system demonstrations, industrial experience reports, reproducibility studies, and benchmark contributions.

Important dates (AoE):

  • Submission deadline: August 29, 2026

  • Notification: September 29, 2026

  • Workshop: NeurIPS 2026, 12-13 December, Paris

Submission instructions and additional details are available on the workshop website. The OpenReview submission portal will be activated shortly and linked from the website.

We would be grateful if you could share this Call for Papers with colleagues, students, and anyone who may be interested.

We look forward to receiving your contributions and to welcoming you at LIGHT during NeurIPS 2026!

Best regards,

Roberta Calegari
University of Bologna

On behalf of the LIGHT Workshop Organizing Committee


Roberta Calegari, PhD
---------------------------------------
Researcher, AI professor 
Department of Computer Science and Engineering  
Alma AI - Alma Mater Research Institute for Human-Centered Artificial Intelligence  
Viale Risorgimento 2, 40126 - Bologna. Italy
---------------------------------------


Paolo Giudici

unread,
Jul 27, 2026, 1:43:17 PM (3 days ago) Jul 27
to AIxIA mailing list

Can We Trust AI Evaluation?


Robustness, Causality, and Risk in Modern AI Assessment
NeurIPS 2026, Sydney, Australia, December 11 or 12, 2026

Overview

Can we trust AI evaluation? Modern AI systems are judged through benchmarks, aggregate scores, and public leaderboards, yet trust in these evaluations is often assumed rather than demonstrated. An evaluation can be precise but measure the wrong construct, stable on a familiar benchmark but brittle on newly collected data, or impressive on a leaderboard while poorly aligned with real-world decisions. Repeated benchmark use, small perturbations, underreported variance, leakage, and contamination can further weaken the evidence behind evaluation claims. The TAE (Trust-AI-Eval): Can We Trust AI Evaluation? workshop treats evaluation itself as an object of study: what is measured, which assumptions connect a protocol to a claim, how uncertainty and failure modes are reported, and when the resulting evidence is strong enough to guide deployment. By bringing together work on robustness, causal and measurement validity, auditing, judge reliability, and deployment risk, the workshop aims to clarify when AI evaluation results deserve trust and how evaluation practices can become more reliable, transparent, and decision-relevant.

We invite submissions on topics including, but not limited to:

  • Uncertainty and robustness: How stable are evaluation conclusions under sampling variation, calibration error, random seeds, data splits, prompts, metrics, evaluator choices, noisy or delayed feedback, tail risk, and worst-case behavior?
  • Benchmark and leaderboard auditing: How do benchmark reuse, contamination, leakage, documentation gaps, lifecycle practice, public incentives, and benchmark-specific optimization affect the trustworthiness of evaluation claims?
  • Black-box auditing: How can AI systems be audited when model internals, training data, or evaluation pipelines are inaccessible, and what behavioral tests, probes, or external evidence can reveal hidden failure modes, contamination, or systematic risk?
  • Measurement and causal validity: What construct is an evaluation protocol intended to measure, what ground truth does it rely on, and what causal, structural, or statistical assumptions connect the protocol to the claim?
  • Stress tests and judge reliability: How should evaluations assess protocol robustness, ambiguous labels, human-, crowd-, and model-judge reliability, and failure modes in evaluation pipelines?
  • Domain coverage and representation: How do imbalances in benchmark suites, such as extensive coverage of coding, mathematics, ethics, and logical reasoning but limited or absent coverage of banking and other regulatory-compliance settings, non-Western cultural contexts, and other underserved domains, affect the validity and generalizability of evaluation claims? How should evaluation portfolios be designed, weighted, and updated to provide representative cross-domain coverage and expose systematic blind spots?
  • Application-domain evaluation: How should evaluation protocols be designed and audited for domain-specific settings such as medicine and healthcare, finance, science, robotics, AI agents, cybersecurity, education, public-sector decision-making, and other high-stakes applications?
  • Deployment risk and governance: When do offline metrics support real-world model selection, safety claims, monitoring, and deployment decisions, and what decision-aware metrics, fairness–accuracy–risk trade-offs, reporting checklists, auditing guidelines, and deployment criteria are needed?

See the Call for Papers for details.

Accepted papers will be presented at the in-person poster session.

Important Dates (Indicative)

Paper submission opens: TBA
Paper submission deadline: August 29, 2026 (AoE)
Review deadline: September 14, 2026 (AoE)
Author notification: September 22, 2026 (AoE)
Final program posted: September 27, 2026
Workshop: December 11 or 12, 2026


Confirmed Speakers & Panelists

Bin Yu
Bin Yu
University of California, Berkeley
Soheil Feizi
Soheil Feizi
University of Maryland
Liming Zhu
Liming Zhu
CSIRO and University of New South Wales
James Bailey
James Bailey
Monash University

Talk titles and panel details will be announced after the final program is confirmed.


Organizers

Hanxun Huang
Hanxun Huang
University of Melbourne
Barbara Tarantino
Barbara Tarantino
University of Pavia
Paolo Giudici
Paolo Giudici
University of Pavia
Xingjun Ma
Xingjun Ma
Fudan University
Eduard Hovy
Eduard Hovy
University of Melbourne
Sarah M. Erfani
Sarah M. Erfani
University of Melbourne

Sponsors

Melbourne ConnectThe University of Melbourne

 •  2026  •  tai-eval.github.io

Theme by beautiful-jekyll



 

Reply all
Reply to author
Forward
0 new messages
Search
Clear search
Close search
Google apps
Main menu