game plan

0 views
Skip to first unread message

Kurt Luther

unread,
May 6, 2013, 5:39:46 PM5/6/13
to crowdcamp-c13
Hi all, I wanted to write up a quick recap of the decisions from our
meeting today, to make sure we're all on the same page and give Paul
an update.


PART I

Part I is a between-subjects experiment with 3 conditions:

- yes/no questions
- yes/no questions, + prompts (character, background, etc.)
- Snowflake method (write 5 sentences)

All 3 conditions begin with a one-sentence story "seed" from the
NYTimes Bestseller list. (This helps even the playing field across
conditions, and provides an abstraction layer.) We will use 5
different seeds. For each seed, we'll have 10 workers trying one of
the 3 conditions, so 30 "pre-stories" per seed, or 150 pre-stories
total. For example:

story seed #1 = 10 workers x 3 conditions
story seed #2 = 10 workers x 3 conditions
etc.

Workers will be paid $0.20 per pre-story.


PART II

In Part II, a different group of workers converts each pre-story into
a longer, 100-word story. Some pre-stories will be formatted as yes/no
questions, others will be 5-sentence snowflakes. Regardless of the
pre-story format, the worker's goal is always to convert it into a
paragraph-length story with roughly 100 words. (Since all conditions
produce the same end product, comparison and evaluation will be
easier.)

Workers are paid $0.20 for each story.


PART III

In Part III, a different group of workers evaluates the 100-word
stories. Each worker is first asked to read the story and summarize it
in one sentence. (This is mainly a gold standard question to ensure
workers read the story, but also produces one-sentence summaries as a
byproduct, cf. machine translation verification.)

Next, the worker evaluates the story, using a 7-pt Likert scale for
each criterion proposed by Pao and Krzysztof (surprise, plot, etc.),
plus a writing quality criterion (to avoid conflating this with other
criteria).

Once the evaluation is complete, the UI reveals the original NYTimes
story "seed" to the worker, and asks her to compare it to the 100-word
story: "how similar is this summary to the long version above?" This
metric helps us understand how faithful the expansions are to the
original.

Workers will be paid $0.15 per evaluation.


Sound about right to everyone?

Kurt


--
Kurt Luther
www.kurtluther.com

Robin

unread,
May 6, 2013, 8:27:07 PM5/6/13
to Kurt Luther, crowdcamp-c13
sounds good to me.

Thanks,
Robin
ROBIN N. BREWER
Human-Centered Computing
University of Maryland, Baltimore County

Kathleen Tuite

unread,
May 6, 2013, 9:01:16 PM5/6/13
to Robin, Kurt Luther, crowdcamp-c13
Sounds good to me, too!

Yotam Gingold

unread,
May 7, 2013, 1:36:17 AM5/7/13
to Kathleen Tuite, Robin, Kurt Luther, crowdcamp-c13
Yes

---
Typed on a tiny keyboard.

Reply all
Reply to author
Forward
0 new messages