Quality of writing- How good is the quality of writing (word choices, uses of grammar and styles)?
Settings
- How original are the settings?
- How believable are the settings?
- How well-developed are the settings?
Characters
- How original are the characters?
- How believable are the characters?
- How well-developed are the characters?
Plot- How original is the plot?
- How believable is the plot?
- How well-developed is the plot?
Page 6, for each story, ask Turker how well does the story match the seed sentence.Story overall- How original is this story overall?
- How believable is this story overall?
- How interesting is this story overall?


I messed around with it and there is a really simple fix: delete line 8,Once you do that, it stops dynamically resizing the Likert scales. The paragraphs don't resize anymore, but that's a small price to pay.Yotam
On Mon, May 27, 2013 at 7:36 PM, Yotam Gingold <yo...@yotamgingold.com> wrote:
The Likert radio buttons still wrap for me (screenshot attached). Can you make the span "non-responsive" or give it a css min-width in "em"s? I messed with it a bit and it might be necessary to also make the span's "float: left" div's to prevent the labels on either end from wrapping. I will try to chat you about this.
Yotam<image.png>
On Mon, May 27, 2013 at 5:48 PM, Kanya(Pao) Siangliulue <ksiang...@gmail.com> wrote:
Hi all,I fixed the UI according to suggestion and add some changes to it. Hopefully this is the final version minus form validation. I will put that up later. Please see http://iis2.seas.harvard.edu/do_eval/eval_stories and ping me if you have any comment.I will be working on setting it ready for mturk this afternoon and will come back to your suggestion before launchingBest,Pao
On Mon, May 27, 2013 at 4:53 AM, Yotam Gingold <yo...@yotamgingold.com> wrote:
Looks nice! Two suggestions:1) Since there is no "Previous" button, can you do some validation to make sure all the questions have been answered before allowing "Next >" to work?2) I think you should set a minimum width for the Likert buttons radio so that they don't wrap; some of the high numbers are wrapping for me and this will give us artificially low numbers.Let me know if I can help.Looking good,Yotam
<Likert-wrap.png>




average
item-test item-rest interitem
Item | Obs Sign correlation correlation covariance alpha
-------------+-----------------------------------------------------------------
set_dev | 964 + 0.8775 0.7103 1.209016 0.6114
set_ori | 964 + 0.8359 0.6266 1.486175 0.7038
set_believe | 964 + 0.7959 0.5392 1.757229 0.7984
-------------+-----------------------------------------------------------------
Test scale | 1.48414 0.7842
-------------------------------------------------------------------------------
average
item-test item-rest interitem
Item | Obs Sign correlation correlation covariance alpha
-------------+-----------------------------------------------------------------
set_dev | 964 + 0.8775 0.7103 1.209016 0.6114
set_ori | 964 + 0.8359 0.6266 1.486175 0.7038
set_believe | 964 + 0.7959 0.5392 1.757229 0.7984
-------------+-----------------------------------------------------------------
Test scale | 1.48414 0.7842
-------------------------------------------------------------------------------
On 29 May 2013, at 17:04, Yotam Gingold <yo...@yotamgingold.com> wrote:
I am suggesting that the confounding factor is when the paragraphs were written.Yotam
On Wed, May 29, 2013 at 10:59 PM, Yotam Gingold <yo...@yotamgingold.com> wrote:
If you look at the Submit Times for the four conditions, you will see that the "blank" stories were mostly written a day later than the others. I am assuming that this is due to the way Amazon randomises HIT assignments when HITs have varying numbers of max_assignments. The blank condition had max_assignments = 10, while the others had max_assignments = 1. I couldn't just submit 10 identical blank tasks instead of 1 task with max_assignments = 10, because a worker might have done the exact same (blank) assignment more than once.
<image.png>Yotam
It seems "overall interesting" and "overall original" may be the same, but they are distinct from "overall believability":
It seems that "setting originality" and "setting well-developed" may be the same, but they are distinct from "setting believability":
The overview you sent on scale construction cautions against assuming that variables are measuring the same thing even if alpha increases unless we have theoretical reasons for doing so: "However, even if the inclusion of age had caused alpha to go up, you wouldn’t want to include it unless you had good theoretical reasons for doing so. (In other words, you don’t just use mindless empiricism when constructing your scales; theory should be guiding you as well.)" Do we have theoretical reasons for doing so?

I just push another set of ratings. Just git pull. All stories now have at least 3 ratings. We probably won't get many more because we ran out of money. This might be a good stopping point to start doing real data analysis. What do people think?The overall significant results are just like what I had earlier. Namely, we still don't get on originality across conditions.
For effect of order of stories, there is only one test with p-value around 0.058 for overall_quality. Otherwise, the order of the stories doesn't effect the outcome.
I ran ICC tests on all individual measures. We don't get any agreement more than 0.5 (I am not surprised). In general, workers don't agree on measure. It would be great if someone can run a test and confirm this fact.I will look into other measures afterwards.
<image.png>

Kruskal-Wallis rank sum testdata: match_sentence by condKruskal-Wallis chi-squared = 25.4934, df = 3, p-value = 1.218e-05
K.
<story oracle analysis 1.pdf>