Neil Harding
unread,Aug 3, 2026, 6:12:10 AM (7 days ago) Aug 3Sign in to reply to author
Sign in to forward
You do not have permission to delete messages in this group
Either email addresses are anonymous for this group or you need the view member email addresses permission to view the original message
to gcd-tech
I've been writing a standalone comic database using the GCD database to automatically identify issues when possible from the filename. It looks for the series name and if there are multiple possible versions it will give you a dropdown, but if it can it will get the data from the GCD database dump. I extract the cover image, but I also take the phash for each cover, and it seems to work well at identifying other copies of the same cover. I have about 200,000 files so once they are all indexed would you like the phashes? I am using Sqlite, Python 3.11, Pony, Bottle, OpenCV, TurboJPG, Scikit Image, Rarlib, 7Z, PDF library, if anyone is interested a copy, and once everything is imported it can also use filehash to identify copies of scans.