Hi Neil,
Your work extracting the covers and calculating perceptual hashes sounds extremely relevant to something we are building.
Simplist identifies comics from user-uploaded photographs. We currently compare full-cover and regional perceptual hashes, including quadrants, so your approach appears closely aligned with ours. Using a hash reference rather than repeatedly retrieving cover images would also fit GCD’s preference to avoid unnecessary image traffic.
Would you be open to discussing access to the hash data for testing? It would be useful to understand:
- how each hash maps back to a GCD issue or cover ID;
- the precise hash algorithm, image preparation and quadrant layout;
- the current coverage and how updates are handled; and
- whether use and caching within a commercial identification application would be permitted.
We would be happy to credit GCD clearly as a data source and work within any reasonable usage requirements.
Thanks,
Rich
Simplist / Bear Maximum