Evaluating plan

1 view
Skip to first unread message

Kanya(Pao) Siangliulue

unread,
May 25, 2013, 1:02:47 PM5/25/13
to crowdcamp-c13
Hi everyone,

Here is our plan for evaluation. We will run it either tonight or tomorrow depending on how fast I code. :)

Each Turker can do 5 HITs (for each seed sentence). For each HIT, a worker will evaluate 4 stories sampled randomly from 4 conditions and they will get paid $0.60. We aim to get at least 5 (striving for 10) workers to evaluate each story.

Details:
Page 1, a Turker sees all four stories in random order. For each story, we ask them to summarize the story into one sentence.

Page 2-5, for each page we show the story and asked the following questions (all are 7-point likert scale),
Quality of writing
- How good is the quality of writing (word choices, uses of grammar and styles)?
Settings
- How original are the settings?
- How believable are the settings?
- How well-developed are the settings?
Characters
- How original are the characters?
- How believable are the characters?
- How well-developed are the characters?
Plot
- How original is the plot?
- How believable is the plot?
- How well-developed is the plot?
Story overall 
- How original is this story overall?
- How believable is this story overall?
- How interesting is this story overall?

Page 6, for each story, ask Turker how well does the story match the seed sentence.

I will be off computer for an hour or so after this but will be online afterwards. If you want to talk, please ping me through gchat or skype.

Best,
Pao

Kanya(Pao) Siangliulue

unread,
May 25, 2013, 4:31:25 PM5/25/13
to crowdcamp-c13
By the way, should we be worried about a chance that turkers in any of previous HITs will be evaluating the stories? Am I too paranoid?

-Pao

Krzysztof Gajos

unread,
May 25, 2013, 7:52:38 PM5/25/13
to Kanya(Pao) Siangliulue, crowdcamp-c13

I hope you are too paranoid, but just in case we should keep track of turker IDs.

Yotam Gingold

unread,
May 25, 2013, 8:09:44 PM5/25/13
to Krzysztof Gajos, Kanya(Pao) Siangliulue, crowdcamp-c13
Yes, I say not to worry about that. If it happens, then we can throw it out, but I really doubt it will.

As for the plan: Do workers complete Pages 1-6 for one story before completing Pages 1-6 for the next story? Or do they complete Page 1 for all 4 stories before moving on to pages 2-5 for all 4 stories, followed by Page 6 for all 4 stories? The former seems "normal" to me, and I wonder what the thinking would be in favor of the latter.

Yotam

Krzysztof Gajos

unread,
May 25, 2013, 9:16:49 PM5/25/13
to Yotam Gingold, Kanya(Pao) Siangliulue, crowdcamp-c13

We started with the "normal" plan, but then moved to the other one: The current plan is to have them read and summarize all four stories first.  Then rate stories one by one and then do page 6 for all of them.  We thought this design would reduce ordering effects and increase consistency in ratings: they won't start rating until they've read all four stories.  Basically, we are trying to increase the power of the cross-condition comparisons.

K.

Krzysztof Gajos

unread,
May 25, 2013, 9:23:01 PM5/25/13
to Kanya Siangliulue, crowdcamp-c13

Pao,

I refined the language of the questions and added anchors:

Quality of writing
- How good is the quality of writing (word choices, grammar and style)?
(1 = very poor; 7 = very good)

Setting
- How well-developed is the setting? 
(1 = not well developed at all; 7 = very well developed)
- How original is the setting?
(1 = not original at all; 7 = very original)
- How believable is the setting?
(1 = not believable at all; 7 = very believable)
 
Characters
- How well-developed are the characters?
- How original are the characters?
- How believable are the characters?
 
Plot
- How well-developed is the plot?
- How original is the plot?
- How believable is the plot?
 
Story overall 
- How original is this story overall?
- How believable is this story overall?
- How interesting is this story overall?
(1 = not interesting at all; 7 = very interesting)

Kanya(Pao) Siangliulue

unread,
May 25, 2013, 9:25:42 PM5/25/13
to Krzysztof Gajos, crowdcamp-c13
Thank you all for the feedback! I will send a link for you guys to look at soon.
-Pao

Kanya(Pao) Siangliulue

unread,
May 26, 2013, 11:33:14 AM5/26/13
to Krzysztof Gajos, crowdcamp-c13
http://458p.localtunnel.com/eval_stories

It is not doing any randomization right now and I will add form validation later. 

-Pao

Krzysztof Gajos

unread,
May 26, 2013, 11:56:12 AM5/26/13
to Kanya(Pao) Siangliulue, crowdcamp-c13

Small suggestions:

- Add numbers to stories (even on the first page so that people can refer to them in their minds as Story 2 or Story 4) and definitely on the subsequent pages so that they know how much progress they've made.

- Change "Characters that appears in the story" to "Characters that appear in the story" (no "s" at the end of "appear")

- On the last page, we want to ask Turkers "how well does the following sentence summarize each of the four stories?"  It is more natural to ask how well a sentence summarizes a story than how well a story matches a summary.  

K.

Kurt Luther

unread,
May 26, 2013, 1:31:00 PM5/26/13
to Krzysztof Gajos, Kanya(Pao) Siangliulue, crowdcamp-c13
Pao, this is great. A few suggestions:

* On the first page, I would strongly encourage you to show only one
paragraph/summary box at a time. In my experience, turkers are really
discouraged by tasks that present a "wall of text", even if they're
paid a lot and the task isn't difficult.

* On the rating page, this too seems a bit overwhelming. It feels like
far more than 13 questions due to all the text on the page. I think it
could help to decrease the space between radio buttons, and increase
the width of the column to the left of the 1s ("very poor", etc.) to
avoid line breaks. The line breaks push the content down the page,
increasing the scrolling (turkers hate scrolling), and increasing the
distance between the paragraph and the ratings.

* On the rating page, I would boldface the key words in each question
(how *original*, how *believable*) etc.

* On the rating page, all of the instances of "very original" on the
far right are missing the 'y'.

* After submitting each rating, is there a way to convey that the
submission was successful before moving onto the next one? The
"autoscroll up" interaction might create confusion if workers think
the system ignored their feedback. They may not immediately realize
there is a new paragraph to be rated.

* The title is "evaluate short short stories" -- joke or typo? :-)

--
Kurt Luther
www.kurtluther.com

Yotam Gingold

unread,
May 26, 2013, 2:27:55 PM5/26/13
to Kurt Luther, Krzysztof Gajos, Kanya(Pao) Siangliulue, crowdcamp-c13
I am guessing that you're working on it now, because the link isn't working for me.

---
Typed on a tiny keyboard.

Kanya(Pao) Siangliulue

unread,
May 26, 2013, 8:55:46 PM5/26/13
to Yotam Gingold, Kurt Luther, Krzysztof Gajos, crowdcamp-c13

I made some changes per suggestion. Still working on dividing the summary page into four small pages and giving feedback when workers submit a form. Sorry for the delay. I will try to push this out tomorrow.

@Kurt. No, short short stories is not a joke. It is another term for flash fiction. http://en.wikipedia.org/wiki/Short_short_story I am not sure which one is more confusing. :-S

Cheers,
Pao

Yotam Gingold

unread,
May 27, 2013, 4:53:11 AM5/27/13
to Kanya(Pao) Siangliulue, Kurt Luther, Krzysztof Gajos, crowdcamp-c13
Looks nice! Two suggestions:

1) Since there is no "Previous" button, can you do some validation to make sure all the questions have been answered before allowing "Next >" to work?

2) I think you should set a minimum width for the Likert buttons radio so that they don't wrap; some of the high numbers are wrapping for me and this will give us artificially low numbers.

Let me know if I can help.

Looking good,
Yotam

Kanya(Pao) Siangliulue

unread,
May 27, 2013, 11:48:51 AM5/27/13
to Yotam Gingold, Kurt Luther, Krzysztof Gajos, crowdcamp-c13
Hi all,

I fixed the UI according to suggestion and add some changes to it. Hopefully this is the final version minus form validation. I will put that up later. Please see http://iis2.seas.harvard.edu/do_eval/eval_stories and ping me if you have any comment. 

I will be working on setting it ready for mturk this afternoon and will come back to your suggestion before launching

Best,
Pao
Likert-wrap.png

Yotam Gingold

unread,
May 27, 2013, 1:36:47 PM5/27/13
to Kanya(Pao) Siangliulue, Yotam Gingold, Kurt Luther, Krzysztof Gajos, crowdcamp-c13
The Likert radio buttons still wrap for me (screenshot attached). Can you make the span "non-responsive" or give it a css min-width in "em"s? I messed with it a bit and it might be necessary to also make the span's "float: left" div's to prevent the labels on either end from wrapping. I will try to chat you about this.

Yotam


Inline image 1
Likert-wrap.png
image.png

Yotam Gingold

unread,
May 27, 2013, 1:40:42 PM5/27/13
to Yotam Gingold, Kanya(Pao) Siangliulue, Kurt Luther, Krzysztof Gajos, crowdcamp-c13
I messed around with it and there is a really simple fix: delete line 8,
<link rel="stylesheet" href= /do_eval/static/css/bootstrap-responsive.min.css >

Once you do that, it stops dynamically resizing the Likert scales. The paragraphs don't resize anymore, but that's a small price to pay.

Yotam

image.png
Likert-wrap.png

Paul André

unread,
May 27, 2013, 10:14:19 PM5/27/13
to Yotam Gingold, Paul André, Kanya(Pao) Siangliulue, Kurt Luther, Krzysztof Gajos, crowdcamp-c13
I have some time tomorrow morning to spend on... outlining with Kurt? Writing notes for a specific section?
I'll chat with Kurt in the morning, or will play it by ear tomorrow!

On 27 May 2013, at 13:40, Yotam Gingold <yo...@yotamgingold.com> wrote:

I messed around with it and there is a really simple fix: delete line 8,
<link rel="stylesheet" href= /do_eval/static/css/bootstrap-responsive.min.css >

Once you do that, it stops dynamically resizing the Likert scales. The paragraphs don't resize anymore, but that's a small price to pay.

Yotam

On Mon, May 27, 2013 at 7:36 PM, Yotam Gingold <yo...@yotamgingold.com> wrote:
The Likert radio buttons still wrap for me (screenshot attached). Can you make the span "non-responsive" or give it a css min-width in "em"s? I messed with it a bit and it might be necessary to also make the span's "float: left" div's to prevent the labels on either end from wrapping. I will try to chat you about this.

Yotam


<image.png>


On Mon, May 27, 2013 at 5:48 PM, Kanya(Pao) Siangliulue <ksiang...@gmail.com> wrote:
Hi all,

I fixed the UI according to suggestion and add some changes to it. Hopefully this is the final version minus form validation. I will put that up later. Please see http://iis2.seas.harvard.edu/do_eval/eval_stories and ping me if you have any comment. 

I will be working on setting it ready for mturk this afternoon and will come back to your suggestion before launching

Best,
Pao
On Mon, May 27, 2013 at 4:53 AM, Yotam Gingold <yo...@yotamgingold.com> wrote:
Looks nice! Two suggestions:

1) Since there is no "Previous" button, can you do some validation to make sure all the questions have been answered before allowing "Next >" to work?

2) I think you should set a minimum width for the Likert buttons radio so that they don't wrap; some of the high numbers are wrapping for me and this will give us artificially low numbers.

Let me know if I can help.

Looking good,
Yotam

<Likert-wrap.png>

Paul André

unread,
May 25, 2013, 5:45:10 PM5/25/13
to Kanya(Pao) Siangliulue, crowdcamp-c13 Project
If you can exclude them, that would be preferable.
Otherwise, you could ask if they did them, and then we just ignore their ratings.

Other than that, evaluation plan seems good! We might want to do a small sample first and see if there's any agreement among Turkers, my experience has shown limited success in evaluating subjective creative artefacts. One way to make it clearer might be to give examples of what we mean by "believable", "well-developed", etc. (Or give positive and negative examples from the stories).

Kanya(Pao) Siangliulue

unread,
May 28, 2013, 11:12:46 AM5/28/13
to Paul André, Yotam Gingold, Kurt Luther, Krzysztof Gajos, crowdcamp-c13
The HITs are up! If you want to test the interface, please be a good worker and actually do the task. :)

See you all later,
Pao

Yotam Gingold

unread,
May 28, 2013, 11:16:08 AM5/28/13
to Paul André, Yotam Gingold, Kanya(Pao) Siangliulue, Kurt Luther, Krzysztof Gajos, crowdcamp-c13
Is anyone working on the writing? I have some time now.

Kurt Luther

unread,
May 28, 2013, 11:44:42 AM5/28/13
to crowdcamp-c13
We've finished our first pass on the outline. We spent a long time
discussing the introduction and various framings, and have included
those notes in the doc, but we think writing that part should wait
until we have results and know what our main takeaways are.

We have lots (too many) of refs for related work, and a basic overview
of our methods. One important TODO is to figure out which stats we'll
run on the evaluation data.

We have four research questions in the Results section; please check
these out and add/remove/modify. This is really the heart of the
paper.

Anyone have some time to take another pass on this and start fleshing
out sections? I don't think I can do much writing in the next few
days.


--
Kurt Luther
www.kurtluther.com


On Tue, May 28, 2013 at 11:23 AM, Paul André
<thethreemo...@gmail.com> wrote:
> Kurt and I are writing in a googledoc, feel free to look it over, but we'll
> be finishing soon, then thinking about where best to go from there.
> https://docs.google.com/document/d/1bZWty6_YlQc-X3eJb7EX-TcR2BlBtqWW9OkxNtqQ7so/edit

Yotam Gingold

unread,
May 28, 2013, 12:17:31 PM5/28/13
to Kurt Luther, crowdcamp-c13
I'll take a pass now.

I set up a template latex document using one of the collaborative online latex sites:

At some point we can copy things over.

Yotam


Yotam Gingold

unread,
May 29, 2013, 6:40:45 AM5/29/13
to Yotam Gingold, Kurt Luther, crowdcamp-c13
I put some more text down for the intro and related work. I copied it over to the writelatex.com document to see how long/how it looks in latex form, but I'm still thinking of the google doc as the "master" document.

How are the turkers doing with their evaluations?

Yotam

Kanya(Pao) Siangliulue

unread,
May 29, 2013, 11:57:25 AM5/29/13
to Yotam Gingold, Kurt Luther, crowdcamp-c13
We have got at least 2 evaluations for all stories. We are still waiting for more results but we also would just do analysis with what we have.  I put the results so far in the git repository. /eval_data/design_oracle_eval_result.csv.

Best,
Pao

Kanya(Pao) Siangliulue

unread,
May 29, 2013, 3:21:33 PM5/29/13
to Yotam Gingold, Kurt Luther, crowdcamp-c13
Please feel free to play with the data. 

Some basic analysis I have so far:
Here is the histogram showing how many ratings each story gets. The x-axis is story_id. (I don't know why Frequency for story 1 is really high. Probably it has something to do with my code.)
Inline image 2

With just raw scoring, there are statistically significant differences across conditions for  overall_believe, plot_believe and set_believe (p-values < 0.05). 

Running 1-way ANOVAs for within subjects shows statistically significance difference across condition for overall_quality, believability and elaboration (well-developed) but NOT originality


The quality of writings are different across conditions (p-value = 0.003102)
Inline image 3



Unsurprisingly prompts has effects on many dependent scores. The prompts have effects on overall quality(average of overall scores), story_believe, plot_believe, ch_believe, ch_developed, set_original, set_believed, set_developed and writing_quality.

I plotted the scores overall_quality and writing_quality below. We can probably treat writing_quality as another factor. I am still figuring out how to do that. If you have some ideas, please share.
Inline image 1
I haven't look at Intraclass correlation yet.

What does this mean? I am still figuring out. Just want to give an update. Feel free to jump in.

-Pao
image.png
image.png
image.png

Kanya(Pao) Siangliulue

unread,
May 29, 2013, 3:25:42 PM5/29/13
to Yotam Gingold, Kurt Luther, crowdcamp-c13
Btw, did I mention that in most (OK all) measures that I am using: overall_quality, believability, originality and elaboration, "blank" condition beats all other cases? 

-Pao
image.png
image.png
image.png

Kurt Luther

unread,
May 29, 2013, 3:33:45 PM5/29/13
to crowdcamp-c13
Hi Pao, thanks so much for getting this started. At the end of our Hangout, I thought Krzysztof found that adding 'writing quality' as a fixed effect caused 'blank' to perform worst and caused the other results to better match our expectations. Did you try this?


--
Kurt Luther
image.png
image.png
image.png

Kanya(Pao) Siangliulue

unread,
May 29, 2013, 3:42:15 PM5/29/13
to Kurt Luther, crowdcamp-c13
Hi Kurt,

Embarrassingly, I did not know how to run two-factor ANOVA for within subject case and how to compare the adjusted means. (Somehow R does not understand the command I put in, even though it seems to work in more than one tutorials.) I am looking for a way to do that right now. If you have a pointer, I can try it out.

Thank you!
Pao
image.png
image.png
image.png

Paul André

unread,
May 29, 2013, 3:48:01 PM5/29/13
to Kurt Luther, Paul André, crowdcamp-c13
Thanks Pao!
Looking at the results now.

Q1: Do Turkers agree on their ratings for stories? Per individual measurement. ICC or some sort of correlation.

Q2: Are our measures conceptually and practically distinct, or are they measuring the same underlying concept? (i.e., do all the setting measures actually correlate highly?). And individual items might suffer from measurement error. This is scale construction, and the way I've done it in the past is calculating Cronbach's alpha. It will allow us to determine if we should/can combine some measurements to create a more reliable and parsimonious scale, and/or if including all measures would lead to problems of multicollinearity in results.

I'm about to go to a talk, but I can look particularly at Q2 soon.

Robin

unread,
May 29, 2013, 4:48:08 PM5/29/13
to Paul André, Kurt Luther, crowdcamp-c13
And I know we had mentioned it yesterday, but does ordering of the stories to rate still affect the ratings?

Thanks,
Robin
ROBIN N. BREWER
Human-Centered Computing
University of Maryland, Baltimore County

Yotam Gingold

unread,
May 29, 2013, 4:59:39 PM5/29/13
to Paul André, Kurt Luther, crowdcamp-c13
If you look at the Submit Times for the four conditions, you will see that the "blank" stories were mostly written a day later than the others. I am assuming that this is due to the way Amazon randomises HIT assignments when HITs have varying numbers of max_assignments. The blank condition had max_assignments = 10, while the others had max_assignments = 1. I couldn't just submit 10 identical blank tasks instead of 1 task with max_assignments = 10, because a worker might have done the exact same (blank) assignment more than once.

Inline image 1

Yotam

image.png

Yotam Gingold

unread,
May 29, 2013, 5:04:09 PM5/29/13
to Yotam Gingold, Paul André, Kurt Luther, crowdcamp-c13
I am suggesting that the confounding factor is when the paragraphs were written.

Yotam
image.png

Paul André

unread,
May 29, 2013, 5:41:48 PM5/29/13
to Yotam Gingold, Paul André, Kurt Luther, crowdcamp-c13
1. dropping anything with NULL in set_dev (1324 obs to 964 observations)
2. scale construction with various subsets (e.g., all character measures, all believability measures), and finally all setting, character, believability, and overall measures.

Including all (twelve) measures we have a Cronbach's alpha of .9485. This is very high, indicating that all measures are essentially measuring the same underlying concept (or that Turkers are unable to differentiate the constructs).

Practically, this means we can argue for combining all measures into a single more reliable scale (unless we want to test specific differences).

Attaching stata output. And here's how to interpret it: http://www3.nd.edu/~rwilliam/stats2/l23.pdf

designoraclealpha.txt

Yotam Gingold

unread,
May 29, 2013, 6:37:04 PM5/29/13
to Paul André, Yotam Gingold, Kurt Luther, crowdcamp-c13
It seems "overall interesting" and "overall original" may be the same, but they are distinct from "overall believability":

                                                            average
                             item-test     item-rest       interitem
Item         |  Obs  Sign   correlation   correlation     covariance      alpha
-------------+-----------------------------------------------------------------
set_dev      |  964    +       0.8775        0.7103        1.209016      0.6114
set_ori      |  964    +       0.8359        0.6266        1.486175      0.7038
set_believe  |  964    +       0.7959        0.5392        1.757229      0.7984
-------------+-----------------------------------------------------------------
Test scale   |                                              1.48414      0.7842
-------------------------------------------------------------------------------


It seems that "setting originality" and "setting well-developed" may be the same, but they are distinct from "setting believability":
                                                            average
                             item-test     item-rest       interitem
Item         |  Obs  Sign   correlation   correlation     covariance      alpha
-------------+-----------------------------------------------------------------
set_dev      |  964    +       0.8775        0.7103        1.209016      0.6114
set_ori      |  964    +       0.8359        0.6266        1.486175      0.7038
set_believe  |  964    +       0.7959        0.5392        1.757229      0.7984
-------------+-----------------------------------------------------------------
Test scale   |                                              1.48414      0.7842
-------------------------------------------------------------------------------

The overview you sent on scale construction cautions against assuming that variables are measuring the same thing even if alpha increases unless we have theoretical reasons for doing so: "However, even if the inclusion of age had caused alpha to go up, you wouldn’t want to include it unless you had good theoretical reasons for doing so. (In other words, you don’t just use mindless empiricism when constructing your scales; theory should be guiding you as well.)" Do we have theoretical reasons for doing so?

Finally, are you able to control for how "well-written" a story was judged to be in the analysis?

Yotam



On 29 May 2013, at 17:04, Yotam Gingold <yo...@yotamgingold.com> wrote:

I am suggesting that the confounding factor is when the paragraphs were written.

Yotam
On Wed, May 29, 2013 at 10:59 PM, Yotam Gingold <yo...@yotamgingold.com> wrote:
If you look at the Submit Times for the four conditions, you will see that the "blank" stories were mostly written a day later than the others. I am assuming that this is due to the way Amazon randomises HIT assignments when HITs have varying numbers of max_assignments. The blank condition had max_assignments = 10, while the others had max_assignments = 1. I couldn't just submit 10 identical blank tasks instead of 1 task with max_assignments = 10, because a worker might have done the exact same (blank) assignment more than once.

<image.png>

Yotam

Paul André

unread,
May 29, 2013, 6:57:59 PM5/29/13
to Yotam Gingold, Paul André, Kurt Luther, crowdcamp-c13
On 29 May 2013, at 18:37, Yotam Gingold <yo...@yotamgingold.com> wrote:

It seems "overall interesting" and "overall original" may be the same, but they are distinct from "overall believability":
It seems that "setting originality" and "setting well-developed" may be the same, but they are distinct from "setting believability":

Both possibly true, but since the alpha goes up a minuscule amount by removing them, we would be safe in combining them.

The overview you sent on scale construction cautions against assuming that variables are measuring the same thing even if alpha increases unless we have theoretical reasons for doing so: "However, even if the inclusion of age had caused alpha to go up, you wouldn’t want to include it unless you had good theoretical reasons for doing so. (In other words, you don’t just use mindless empiricism when constructing your scales; theory should be guiding you as well.)" Do we have theoretical reasons for doing so?

I think there's a an argument to be made that they are testing slightly different variations on a core question: "how good is this story?". I spent *months* trying to get turkers (and odeskers, and creative writing majors) to rate limericks on different dimensions, but they're all so highly correlated (humor, flow, coherence, rhyme, overall quality) that we just ended up always combining measures.

If we had specific (theory or intuition) about testing the measures separately, which I think we are interested in, we can always test them. However, to get a more reliable measure, we can also combine them if we want. (This also doesn't take into account whether Turkers even agree for specific stories, which we can judge as more ratings come in. Something which I am more concerned about ;p)

Kanya(Pao) Siangliulue

unread,
May 29, 2013, 8:19:13 PM5/29/13
to Paul André, Yotam Gingold, Kurt Luther, crowdcamp-c13
I just push another set of ratings. Just git pull. All stories now have at least 3 ratings. We probably won't get many more because we ran out of money. This might be a good stopping point to start doing real data analysis. What do people think?

The overall significant results are just like what I had earlier. Namely, we still don't get on originality across conditions.

For effect of order of stories, there is only one test with p-value around 0.058 for overall_quality. Otherwise, the order of the stories doesn't effect the outcome.

I ran ICC tests on all individual measures. We don't get any agreement more than 0.5 (I am not surprised). In general, workers don't agree on measure. It would be great if someone can run a test and confirm this fact.

I will look into other measures afterwards.

Inline image 1
image.png

Krzysztof Gajos

unread,
May 29, 2013, 11:10:20 PM5/29/13
to Kanya(Pao) Siangliulue, Paul André, Yotam Gingold, Kurt Luther, crowdcamp-c13

About including writing quality in the analysis: you won't be able to do it for within-subjects analysis because writing quality differed per story and not per participant.  You *can* use writing quality as a covariate if you instead treat turker id as a random factor: this gives you some (but not all) benefits of within subjects analysis, but allows you to include per-story covariates like writing quality.

I will run analyses on my end in a moment.  Greetings from Pittsburgh.

K.

On May 29, 2013, at 8:19 PM, "Kanya(Pao) Siangliulue" <ksiang...@gmail.com> wrote:

I just push another set of ratings. Just git pull. All stories now have at least 3 ratings. We probably won't get many more because we ran out of money. This might be a good stopping point to start doing real data analysis. What do people think?

The overall significant results are just like what I had earlier. Namely, we still don't get on originality across conditions.

For effect of order of stories, there is only one test with p-value around 0.058 for overall_quality. Otherwise, the order of the stories doesn't effect the outcome.

I ran ICC tests on all individual measures. We don't get any agreement more than 0.5 (I am not surprised). In general, workers don't agree on measure. It would be great if someone can run a test and confirm this fact.

I will look into other measures afterwards.

<image.png>

Krzysztof Gajos

unread,
May 29, 2013, 11:37:07 PM5/29/13
to Kanya(Pao) Siangliulue, Paul André, Yotam Gingold, Kurt Luther, crowdcamp-c13

I am attaching results of two analyses.  In both turker id was modeled as a random factor.  In the first, the fixed factors were: prompt, condition, order in which raters saw the stories, writing quality and the interaction between prompt and condition.  In the second analysis, I added the overall development score (how well the different parts of the story were developed) as a covariate.  The rationale here is that the depth of development is part of the craft rather than part of the raw divergent thinking so I wanted to control for it.  Here are some highlights:

In analysis 1:

-- no effect of condition on overall quality or originality
-- condition significantly impacts believability (blank best, qs worst)
-- big effect of condition on how well developed the stories were: snowflake best, qs worst

In analysis 2 (with level of development as a covariate):

-- significant effect of condition on originality! qstructured best (!), then qs, then blank, then snowflake
-- significant effect of condition on overall quality: blank > qstructured > qs > snowflake
-- significant effect of condition on believability: blank > qstructured > qs > snowflake

This said, the magnitudes of the differences are pretty small.  The output of the analyses is attached.  The Least Squares Means reported are the estimates of the true impacts of the different factors as produced by the regression models (so they are not simple arithmetic means of the scores).

K.

P.S. Google groups objected to the size of the attachments so I have to send them in two separate emails.  Below is analysis 1:
story oracle analysis 1.pdf

Krzysztof Gajos

unread,
May 29, 2013, 11:38:01 PM5/29/13
to Kanya(Pao) Siangliulue, Paul André, Yotam Gingold, Kurt Luther, crowdcamp-c13

And now analysis 2:

story oracle analysis 2.pdf

Kanya(Pao) Siangliulue

unread,
May 30, 2013, 1:05:03 AM5/30/13
to Krzysztof Gajos, Paul André, Yotam Gingold, Kurt Luther, crowdcamp-c13
Interesting results. I took a peek at how well a story matched its original sentence. See raw means below. The whiskers are for 95% confidence intervals.
Inline image 1
A Kruskal-Wallis test shows a statistically significant difference across condition. 

Kruskal-Wallis rank sum test

data:  match_sentence by cond 
Kruskal-Wallis chi-squared = 25.4934, df = 3, p-value = 1.218e-05

It is not strange for a story to diverge from the original sentence more if we intervene with yes-no question. I am surprised by the snowflake condition because I expected this case to be the one that sticks closest to the original sentence. It is still a stretch to say that by using yes-no question we are making people diverge from original prompts (and the generated stories are not more interesting anyway).

Question:
What are our main conclusions? Do we have something for the paper? I can write most of tomorrow and Friday.

Cheers,
Pao



K.


image.png

Paul André

unread,
May 30, 2013, 10:21:47 AM5/30/13
to Krzysztof Gajos, Paul André, Kanya(Pao) Siangliulue, Yotam Gingold, Kurt Luther, crowdcamp-c13
Thanks Krzysztof.

So one analysis has: no effect of condition on overall quality, effect on believability with blank best, and effect on how well developed with snowflake best. This doesn't seem particularly interesting: a structured method (snowflake) is likely to produce better developed stories.

The other analysis has, which I believe the rationale for, has: effect on originality with qstructured best; and effect on quality and believability with blank best. This is unfortunate. The originality finding is interesting, but on its own quite weak.

Add in that Turker's didn't seem to agree on story rating (ICC=.5), I'm not sure we have a contribution or story at this point. It's a shame not to see all this work result in something, but given these results and how close the deadline is I am not sure we have a story. 

Does anyone else see an angle here? Or a possible revision for a different venue? (HCOMP WiP is July).

cheers,
Paul

<story oracle analysis 1.pdf>

Yotam Gingold

unread,
May 30, 2013, 5:48:13 PM5/30/13
to Paul André, Krzysztof Gajos, Kanya(Pao) Siangliulue, Yotam Gingold, Kurt Luther, crowdcamp-c13
We only had one shot at getting good data and a good evaluation, and it seems like we don't quite have a result worth sharing yet. To think about next steps: We have a nice trove of collected data, and we may be able to iterate the evaluation (keeping the data). We may also want to collect more data.

I will say that myself, I didn't have time to sit down and read the story facts or the paragraphs that were evaluated. I wonder if that will reveal something. I plan to look at them. I wonder if we limited the range of stories (originality) by starting all conditions from the same five sentences; we could instead use our four conditions to generate from scratch, and then distill that back down into a single sentence (our abstraction barrier), and compare those (forced choice most creative evaluation?).

Yotam

Reply all
Reply to author
Forward
0 new messages