|
Caedral
AI infrastructure and retrieval systems
|
Hello HotpotQA Team,
My name is Leonardo Turque, founder of
Caedral.
We develop AI infrastructure and CPU-efficient embedding and reranking systems.
We are evaluating HotpotQA as a training source for proprietary English retrieval
models. We understand that the dataset and the processed Wikipedia corpus are
distributed under Creative Commons Attribution-ShareAlike 4.0.
We would like written clarification regarding commercial model training,
private model weights and hosted inference.
Requested clarification
-
Whether HotpotQA may be used to train proprietary commercial embedding
and reranking models.
-
Whether resulting model weights may remain private and be served
commercially through hosted APIs.
-
Whether filtering, deduplication, transformation, hard-negative generation,
teacher scoring and combination with other datasets are permitted.
-
Whether ShareAlike applies to model weights and internal training artifacts,
or only to redistributed adaptations of the dataset.
-
Which attribution language should appear in model cards, dataset cards
and public documentation.
Caedral would not redistribute the raw dataset. Access would be restricted to the
training environment, with provenance, versioning, security and deletion controls.
Could you please confirm whether the uses above are already permitted by the
current license, or direct us to the appropriate rights holder for a separate
written authorization?
Thank you for your time.
Leonardo Turque
Founder, Caedral
https://caedral.com
sup...@caedral.com
@trycaedral
Caedral · AI infrastructure, embeddings and reranking
https://caedral.com
|