Thank you for sharing! It is an interesting project. Reading your about page, I think your "gap" framing can be a litte misleading. There are a couple of reasons for why I think this:1) It's true that SE produces an average rate of 127 books per year, but the current rate is more like 200 or more per year. You can confirm this by looking at the bulk downloads page. This rate has been increasing over time as more and more contributors join the project. So the "gap" between SE and PG is being closed at a faster and faster rate.2) Only a portion of the PG catalog is relevant to SE. We exclude, for example, foreign-language texts, and outdated instructional books like Glue, Gelatine, Animal Charcoal, Phosphorous, Cements, Pastes and Mucilages (really interesting stuff). We pick a single edition/translation for each SE book, but PG may have multiple editions and translations of the same book. Also in some cases, like short fiction compilations and books that were originally published in multiple volumes, one SE book corresponds to multiple PG books.3) The same argument applies to the public domain as a whole. How many of the 30-40 million books in the public domain actually fall under the SE collection policy and are interesting enough that a contemporary reader would want to pick it up? If it's something around one percent (not an unreasonable estimate in my opinion), the number comes down to 300,000 or 400,000.4) There is also something of a "less is more" element with the SE collection. The fact that people work on projects end-to-end means that SE becomes a sort of collectively curated catalog of interesting books in the public domain. I think part of the charm of our project is that our collection is more of a conglomeration of eclectic interests than a top-down effort to curate a canon.On the practical side, I think it may be difficult for your site to get the traction you would like. My general sense is that this space relies a lot on both first-mover advantage and word of mouth. You may have a hard time getting readers to buy in to what you are doing, especially if part of your sell is that your project is highly automated. Also, as I mentioned in some of our emails off of this list, I think that a more fruitful direction for your transcription framework would be to build a tool that can transcribe a book and convert it into a basic Markdown or HTML format. In my opinion, the work of getting a usable transcription is much more useful and worthy of automation than the PG -> SE reformatting steps.
You received this message because you are subscribed to a topic in the Google Groups "Standard Ebooks" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/standardebooks/oMPyQCotQh8/unsubscribe.
To unsubscribe from this group and all its topics, send an email to standardebook...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/standardebooks/59db8ac6-d693-4084-9b7d-9d6fe701f670%40standardebooks.org.
The bigger picture is that there are an estimated 30-40 million public domain works. SE has produced ~1,500 titles. Even at current pace, closing that gap with current methods is a multi-century project, if ever. This is my own answer to the question of how to bridge the gap.
I want to share an automated production pipeline I've been creating over the last 3 months that takes a PG transcription from raw text to a finished, well-typeset EPUB — in 1-2 days per book. I aimed for 90-95% SE quality...you can judge for yourself at the link below.
Based on what I've seen of the rate of book creation here, 1 hour for a book is not a claim I can believe without further proof. I've been seeing 3 days for a short work from an expert, all the way up to 6 months for someone new. That is totally fine, and I commend these folks.
Without having read the thread in the greatest detail it does seem that Travis's aims are very different from SE's. So I'm not sure whether this will interest you or not, Travis, but another reason the gap metaphor is inaccurate is that SE hosts work that PG not only lacks but, given its rules and aims, is always going to lack. I am thinking of the omnibus collections in particular. SE has sometimes ended up hosting the first and only complete collection of an author's PD pieces within a particular genre (poetry, plays, short fiction, short works). The content of such collections is the product of extensive bibliographical research, not least because the SE rule is that a collection should contain everything PD that can be found -- all the poems/plays/stories/essays, which the author has usually published in a variety of independent journals, some of them quite obscure and not always available on Internet Archive/HathiTrust. Editors like Weijia have sometimes transcribed pieces from print sources that had to be ordered from/accessed through libraries, because they hadn't been digitised (and perhaps never will be). The Yeats poetry collection produced by Christopher Hapka is a good example of this. The editors can probably remember many other examples.
PG will publish volumes of journals and collections of works, but they will not create collections like these of SE's from an author's individually published pieces, unless those have already been collected and published in an independent anthology that has then become PD itself. This isn't at all a failing of PG given their aims, but it shows that SE's distinct aims lead to valuable work that isn't currently found at PG and never will be unless they change their aims. The rate of ebook release alone tells us nothing about this.
This isn't to repeat Weijia's fourth (good) point about careful curation of a corpus, but is to say further that the SE corpus does contain collections of a kind, and with a value, that can't currently be found at PG, or anywhere else in the public domain, and sometimes not anywhere else at all.
--
You received this message because you are subscribed to the Google Groups "Standard Ebooks" group.
To unsubscribe from this group and stop receiving emails from it, send an email to standardebook...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/standardebooks/625f84f1-5b80-4030-b854-63edaba930a2n%40googlegroups.com.