Project post: transorthogonal-linguistics

287 views
Skip to first unread message

Travis Hoppe

unread,
Aug 25, 2015, 10:19:13 AM8/25/15
to gensim
Hello gensim users!

A few months ago I made a project that utilized gensim's word2vec for a local Hack and Tell. I'm posting this here to share my experience. Essentially it uses two W2V vectors as start and end points and tries to find the words _between_ the two points -- a semantic passage if you will. The project writeup is here:


with a live demo:


(it's Heroku free tier so give it a moment to spin up). 

Note to aspiring developers/users. I found that the full W2V vectors given by gensim were large and unwieldy (especially for a live deployment). I took most of the data out myself and saved them to numpy arrays -- it would nice if there was an option to do this automatically. 

Thanks for the awesome software!
Travis Hoppe

Shayne Miel

unread,
Aug 25, 2015, 2:11:57 PM8/25/15
to gen...@googlegroups.com
Very cool!!! Thanks for sharing.

--
You received this message because you are subscribed to the Google Groups "gensim" group.
To unsubscribe from this group and stop receiving emails from it, send an email to gensim+un...@googlegroups.com.
For more options, visit https://groups.google.com/d/optout.

Shayne Miel

unread,
Aug 25, 2015, 2:26:12 PM8/25/15
to gen...@googlegroups.com, emay...@turnitin.com
I shared this with my research group and one of my co-workers ran sociology -> mathematics. Look how well it replicates the xkcd comic "Purity" (https://xkcd.com/435/)!


purity.png
Screen Shot 2015-08-25 at 2.17.45 PM.png

hoppe

unread,
Aug 25, 2015, 2:45:16 PM8/25/15
to gen...@googlegroups.com
That's amazing! Thanks for finding that. Please post any more good ones you find and I'll add them as examples to the repo!

You received this message because you are subscribed to a topic in the Google Groups "gensim" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/gensim/a8PF8CInRKk/unsubscribe.
To unsubscribe from this group and all its topics, send an email to gensim+un...@googlegroups.com.

Shayne Miel

unread,
Aug 25, 2015, 3:17:57 PM8/25/15
to gen...@googlegroups.com
Well, "hello" -> "goodbye" tells an amazing and NSFW story. :)

Gordon Mohr

unread,
Aug 25, 2015, 3:28:45 PM8/25/15
to gensim
Very interesting!

Do I understand correctly that a word *could* be returned as being 'near the path' even if its nearest-point is slightly outside the t=0 to t=1 range (beyond each endpoint), as long as it's still in the top-25 closest words? (If not, any chance you could point me at the part of the distance-calc that precludes that?)

It might be interesting to report (or plot) the example words' individual distances from the straight-line (chord?). 

If my intuition from 2D/3D is correct (and of course it might be very wrong in high dimensions!), the "cheat a little" approximation should always yield the correct relative t values (ordering along the line), but will exaggerate the distance-from-the-arc more as t is closest to 0.5. (That is, where the chord-approximation is most different from the hypersphere arc.) So the top-N clipping may favor words nearer the endpoints... but perhaps a distance-corrective factor as a function of t would be straightforward?

That correlation with the XKCD that miel.shayne forwards is something else! I suspect other wry but more-complicated navigations should also be possible, like: "To get to [famous person X], start at [famous person A], go halfway to [famous person B], turn towards [personality adjective C], and stop just past [famous person Y]."

Regarding: "I found that the full W2V vectors given by gensim were large and unwieldy (especially for a live deployment). I took most of the data out myself and saved them to numpy arrays -- it would nice if there was an option to do this automatically."

The word vectors in the Word2Vec instance are already a numpy array (model.syn0) – so unsure what you mean here, other than simply taking that array directly when done training. (Or maybe, discarding some of its rows or reducing the float resolution to save space?) Is it that you'd like an option to discard all the other state (making the model read-only) when no longer needed for more training?

- Gordon

hoppe

unread,
Aug 25, 2015, 3:46:31 PM8/25/15
to gen...@googlegroups.com
> It might be interesting to report (or plot) the example words' individual distances from the straight-line (chord?). 

Yes it might be, but I wasted a lot of time trying to come up with an intuitive visual for this. Everything I tried ended up being messier than the picture in my head.

> Do I understand correctly that a word *could* be returned as being 'near the path' even if its nearest-point is slightly outside the t=0 to t=1 range (beyond each endpoint), as long as it's still in the top-25 closest words? (If not, any chance you could point me at the part of the distance-calc that precludes that?)

There is nothing that precludes that at all. The reason why it is "transorthogonal" if you haven't caught that is that the calculation of the nearest point to a line is simple in a Cartesian coordinates, even if it is R^300. In fact, finding the closest point to a line when given a _set_ of points is a linear algebra operation -- this was the key insight that allowed me to find it points quick enough for a little web demo.

I never quite figured out a proper metric though that was good for all paths, so I found that in the end I would simply show the top 25 words. The vertical distance you see from one word to a another is representative of the time t=0 [word A] to t=1 [word B].

> If my intuition from 2D/3D is correct (and of course it might be very wrong in high dimensions!), the "cheat a little" approximation should always yield the correct relative t values (ordering along the line), but will exaggerate the distance-from-the-arc more as t is closest to 0.5. (That is, where the chord-approximation is most different from the hypersphere arc.) So the top-N clipping may favor words nearer the endpoints... but perhaps a distance-corrective factor as a function of t would be straightforward?

Yes this is true, and you can do a taylor series expansion to find out an approximate error and when the hyperchord approximation will break down. It can be worse that 0.5 though -- the two words can sit on opposite ends of the hypersphere! Yes, it will favor words near the endpoints -- so I actually did try your suggestion by breaking it up into little short line segments. If you're reading this far down there is a hidden button on the app. Press the [*] in the title to switch to this mode (FYI this was my first flask app and everyone is sharing the same instance!). IMHO it doesn't work as well for reasons I don't fully understand.

> That correlation with the XKCD that miel.shayne forwards is something else! I suspect other wry but more-complicated navigations should also be possible, like: "To get to [famous person X], start at [famous person A], go halfway to [famous person B], turn towards [personality adjective C], and stop just past [famous person Y]."

Could be very cool, but I did away with that information early on when I decided to stick with one word tokens (keeping case, e.g. god != God). Clinton is ambiguous as to which one for example. Multi-word tokenization is a topic of itself, and I need to look into that a bit more.

> The word vectors in the Word2Vec instance are already a numpy array (model.syn0) – so unsure what you mean here, other than simply taking that array directly when done training. (Or maybe, discarding some of its rows or reducing the float resolution to save space?) Is it that you'd like an option to discard all the other state (making the model read-only) when no longer needed for more training?

Yes the vectors are numpy arrays, and this is what I saved in the end. When you save them from gensim however, I think it pickles them -- keeping the large class and methods information with it. This additionally required loading the library each time for post-analysis so I stripped out the numpy part directly.

--

Gordon Mohr

unread,
Aug 25, 2015, 7:16:52 PM8/25/15
to gensim
On Tuesday, August 25, 2015 at 12:46:31 PM UTC-7, Travis Hoppe wrote:
> It might be interesting to report (or plot) the example words' individual distances from the straight-line (chord?). 

Yes it might be, but I wasted a lot of time trying to come up with an intuitive visual for this. Everything I tried ended up being messier than the picture in my head.

My first thought was something with d3 & a force-directed graph. While locking a straight line between the endpoints, each 'transorth' word would have an strong edge to their projection point ('t'?), or perhaps to each of the endpoints. (They might have weaker edges to other 'transorths'.) This might even be survive to sets of 3, 4, 5 endpoint-words. 

Another variant that could offer nifty visuals might be to bias the path towards some chosen third concept. ("A to B, bulging towards C.") In the extreme this might yield subway-map like graphs, or 'concept curves' (still really line-segment approximations) that can be adjusted with handles a bit like the bezier handles of vector-drawing programs.
 
> Do I understand correctly that a word *could* be returned as being 'near the path' even if its nearest-point is slightly outside the t=0 to t=1 range (beyond each endpoint), as long as it's still in the top-25 closest words? (If not, any chance you could point me at the part of the distance-calc that precludes that?)

There is nothing that precludes that at all. The reason why it is "transorthogonal" if you haven't caught that is that the calculation of the nearest point to a line is simple in a Cartesian coordinates, even if it is R^300. In fact, finding the closest point to a line when given a _set_ of points is a linear algebra operation -- this was the key insight that allowed me to find it points quick enough for a little web demo.

I never quite figured out a proper metric though that was good for all paths, so I found that in the end I would simply show the top 25 words. The vertical distance you see from one word to a another is representative of the time t=0 [word A] to t=1 [word B].

And when multiple words are on one line, is that because they all fall within the same contiguous bucket of t ranges? 
 
> If my intuition from 2D/3D is correct (and of course it might be very wrong in high dimensions!), the "cheat a little" approximation should always yield the correct relative t values (ordering along the line), but will exaggerate the distance-from-the-arc more as t is closest to 0.5. (That is, where the chord-approximation is most different from the hypersphere arc.) So the top-N clipping may favor words nearer the endpoints... but perhaps a distance-corrective factor as a function of t would be straightforward?

Yes this is true, and you can do a taylor series expansion to find out an approximate error and when the hyperchord approximation will break down. It can be worse that 0.5 though -- the two words can sit on opposite ends of the hypersphere! Yes, it will favor words near the endpoints -- so I actually did try your suggestion by breaking it up into little short line segments. If you're reading this far down there is a hidden button on the app. Press the [*] in the title to switch to this mode (FYI this was my first flask app and everyone is sharing the same instance!). IMHO it doesn't work as well for reasons I don't fully understand.

Aha, I'd found that toggle, was wondering what 'SLERP' variant meant. In my probes, those results seem (sometimes) more interesting – I suppose whether they're truly better or worse could depend on taste or ultimate goals. 

(For easier comparison it'd be helpful to either show both default/SLERP on same page, or retain the seed word pair when toggling.)
 
> That correlation with the XKCD that miel.shayne forwards is something else! I suspect other wry but more-complicated navigations should also be possible, like: "To get to [famous person X], start at [famous person A], go halfway to [famous person B], turn towards [personality adjective C], and stop just past [famous person Y]."

Could be very cool, but I did away with that information early on when I decided to stick with one word tokens (keeping case, e.g. god != God). Clinton is ambiguous as to which one for example. Multi-word tokenization is a topic of itself, and I need to look into that a bit more.

You may want to check out "Document Embedding With Paragraph Vectors": http://arxiv.org/abs/1507.07998

In a way, it suggests a potential alternative to multi-word-tokenization or concept-extraction: full-articles (by their title) as discrete entities. Even if text analysis never learns 'Hillary Clinton' (and all variants) as one entity, there's still the vector for the article "Hillary Clinton", which is in the same "space" as words and (plausibly) useful almost as if it were a superword. 

I've run a 300d DBOW+words job on a May 2015 Wikipedia dump that might be useful for this – but it's many gigabytes in size. (The 1 million word vocabulary array is about 1.2GB; the 4.8 million full-article vectors are about 5.5GB.)

> The word vectors in the Word2Vec instance are already a numpy array (model.syn0) – so unsure what you mean here, other than simply taking that array directly when done training. (Or maybe, discarding some of its rows or reducing the float resolution to save space?) Is it that you'd like an option to discard all the other state (making the model read-only) when no longer needed for more training?

Yes the vectors are numpy arrays, and this is what I saved in the end. When you save them from gensim however, I think it pickles them -- keeping the large class and methods information with it. This additionally required loading the library each time for post-analysis so I stripped out the numpy part directly.

I see. The Word2Vec `save()` method uses numpy's own save routine to store large arrays as separate files, so I *think* if you merely nulled out the parts of the model that were no longer needed, gensim's own saved version would be nearly as compact as any other approach. (The base model classes might need to be made more tolerant of such a stripped state.)

- Gordon

Radim Řehůřek

unread,
Aug 25, 2015, 9:48:01 PM8/25/15
to gensim
Haha, nice app Travis!

Regarding storing objects: no, pickle doesn't save any class nor method information. It only stores the state.

Regarding removing unneeded parts: to make a trained word2vec model smaller and read-only, `model.init_sims(replace=True)`. There's still a little overhead compared to "storing the matrix only", because this keeps the vocabulary mapping in, but probably worth it for having a model with sane API access to the vectors (which matrix row is which word?).

Best,
Radim
Reply all
Reply to author
Forward
0 new messages