You will read research papers, one at a time. Each paper introduces a
test for AI systems. Your job is to answer a fixed
list of questions about what the paper itself says — for
example: does the paper say what ability its test is supposed to
measure? If yes, copy the exact sentence where it says so.
One note on wording. The menus and buttons on this page say
paper, because you work through one paper at a time. Some of
the questions say benchmark instead. That is the usual research
word for the test a paper introduces. So when a question asks
about “the benchmark”, it is asking about the test
described in the paper you are reading right now.
Two ground rules:
- Only the paper counts. Answer from what is written in the
PDF on the left — not from anything you may already know or
assume. If the paper doesn't say it, the answer is “no” or
“not stated”, even if it seems obvious.
- There is no right answer to match. Nobody can see your
answers and you can't see anyone else's. Later, answers are
compared and differences discussed — honest disagreement is
a normal and useful part of the process. Never guess what
“they” want; say what you see in the paper.