Commercial model training clarification for HotpotQA — Caedral

2 views
Skip to first unread message

sup...@caedral.com

unread,
Jul 31, 2026, 2:33:41 PM (9 days ago) Jul 31
to hotp...@googlegroups.com

Caedral

AI infrastructure and retrieval systems

Hello HotpotQA Team,

My name is Leonardo Turque, founder of Caedral. We develop AI infrastructure and CPU-efficient embedding and reranking systems.

We are evaluating HotpotQA as a training source for proprietary English retrieval models. We understand that the dataset and the processed Wikipedia corpus are distributed under Creative Commons Attribution-ShareAlike 4.0.

We would like written clarification regarding commercial model training, private model weights and hosted inference.

Requested clarification

  • Whether HotpotQA may be used to train proprietary commercial embedding and reranking models.
  • Whether resulting model weights may remain private and be served commercially through hosted APIs.
  • Whether filtering, deduplication, transformation, hard-negative generation, teacher scoring and combination with other datasets are permitted.
  • Whether ShareAlike applies to model weights and internal training artifacts, or only to redistributed adaptations of the dataset.
  • Which attribution language should appear in model cards, dataset cards and public documentation.

Caedral would not redistribute the raw dataset. Access would be restricted to the training environment, with provenance, versioning, security and deletion controls.

Could you please confirm whether the uses above are already permitted by the current license, or direct us to the appropriate rights holder for a separate written authorization?

Thank you for your time.

Leonardo Turque
Founder, Caedral
https://caedral.com
sup...@caedral.com
@trycaedral


Caedral · AI infrastructure, embeddings and reranking

https://caedral.com

Reply all
Reply to author
Forward
0 new messages