--
You received this message because you are subscribed to the Google Groups "gensim" group.
To unsubscribe from this group and stop receiving emails from it, send an email to gensim+un...@googlegroups.com.
For more options, visit https://groups.google.com/d/optout.
You received this message because you are subscribed to a topic in the Google Groups "gensim" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/gensim/a8PF8CInRKk/unsubscribe.
To unsubscribe from this group and all its topics, send an email to gensim+un...@googlegroups.com.
--
> It might be interesting to report (or plot) the example words' individual distances from the straight-line (chord?).Yes it might be, but I wasted a lot of time trying to come up with an intuitive visual for this. Everything I tried ended up being messier than the picture in my head.
> Do I understand correctly that a word *could* be returned as being 'near the path' even if its nearest-point is slightly outside the t=0 to t=1 range (beyond each endpoint), as long as it's still in the top-25 closest words? (If not, any chance you could point me at the part of the distance-calc that precludes that?)There is nothing that precludes that at all. The reason why it is "transorthogonal" if you haven't caught that is that the calculation of the nearest point to a line is simple in a Cartesian coordinates, even if it is R^300. In fact, finding the closest point to a line when given a _set_ of points is a linear algebra operation -- this was the key insight that allowed me to find it points quick enough for a little web demo.I never quite figured out a proper metric though that was good for all paths, so I found that in the end I would simply show the top 25 words. The vertical distance you see from one word to a another is representative of the time t=0 [word A] to t=1 [word B].
> If my intuition from 2D/3D is correct (and of course it might be very wrong in high dimensions!), the "cheat a little" approximation should always yield the correct relative t values (ordering along the line), but will exaggerate the distance-from-the-arc more as t is closest to 0.5. (That is, where the chord-approximation is most different from the hypersphere arc.) So the top-N clipping may favor words nearer the endpoints... but perhaps a distance-corrective factor as a function of t would be straightforward?Yes this is true, and you can do a taylor series expansion to find out an approximate error and when the hyperchord approximation will break down. It can be worse that 0.5 though -- the two words can sit on opposite ends of the hypersphere! Yes, it will favor words near the endpoints -- so I actually did try your suggestion by breaking it up into little short line segments. If you're reading this far down there is a hidden button on the app. Press the [*] in the title to switch to this mode (FYI this was my first flask app and everyone is sharing the same instance!). IMHO it doesn't work as well for reasons I don't fully understand.
> That correlation with the XKCD that miel.shayne forwards is something else! I suspect other wry but more-complicated navigations should also be possible, like: "To get to [famous person X], start at [famous person A], go halfway to [famous person B], turn towards [personality adjective C], and stop just past [famous person Y]."Could be very cool, but I did away with that information early on when I decided to stick with one word tokens (keeping case, e.g. god != God). Clinton is ambiguous as to which one for example. Multi-word tokenization is a topic of itself, and I need to look into that a bit more.
> The word vectors in the Word2Vec instance are already a numpy array (model.syn0) – so unsure what you mean here, other than simply taking that array directly when done training. (Or maybe, discarding some of its rows or reducing the float resolution to save space?) Is it that you'd like an option to discard all the other state (making the model read-only) when no longer needed for more training?Yes the vectors are numpy arrays, and this is what I saved in the end. When you save them from gensim however, I think it pickles them -- keeping the large class and methods information with it. This additionally required loading the library each time for post-analysis so I stripped out the numpy part directly.