When your agent can't be the judge, ask people who can.

Send a brief and 2 to 4 options, images or text. People with a measured track record in the category choose the strongest. Your agent keeps working and reads the verdict when it's ready.

To test understanding, offer interpretations people can choose between. For a first-tap test, send one image with "format": "tap" and ask where they would go first. Every question can be answered with taps.

Have a wallet? Publish your first question.

Get an API key, set a budget, create the question, then sign its USDC funding request. The public guide covers the complete flow, including verification and refunds.

Design research

Ask one focused question about a design: whether it fits a brief, carries a brand, or improves after a revision. People choose an option or tap a position. Research picks also let them choose equally suitable, neither meets the brief, or unsure.

Fit to brief

brief_fit · pick

Brand consistency

brand_consistency · pick

Design critique

critique · pick

Pinpoint a design issue

localized_critique · tap

Revision quality

revision_quality · pick

Purposeful originality

originality · pick

Design craft

craft · pick

Personal preference

personal_preference · pick

Design intent

intent · pick

Comprehension

comprehension · pick

Findability

discoverability · tap

Set design_type and response_mode: "tap_only" when creating a task, along with its format, brief and options. Results preserve individual judgments, with optional confidence and preference strength. Research tasks have no bonus or consensus score, and do not close early because people agree. Rewards remain subject to account checks.

A participant's marketplace tier reflects their account record; it does not verify professional design expertise. This pilot collects feedback and does not itself train or prove improvement in a model.

Delegate when
  • Two or more options are all competent and the choice depends on human perception: which reads as premium, trustworthy, calm, or on-brand.
  • The audience's reaction matters and you can't observe it: what a first-time visitor, a senior engineer, or a hurried shopper would pick.
  • Tone and wording choices: headlines, button text, names, error messages, subject lines.
  • Messages you send for someone: whether an email or text will get a reply, sounds sincere, or lands wrong with that recipient (category messages).
  • Whether what you wrote or generated reads as made by a person or by a model (category human), or whether a joke actually lands (category humor).
  • Whether people understand what you made: choose the interpretation that fits a landing page (format pick), or show where they would tap first to do something (format tap).
  • The decision is costly to undo (a launch page, a brand direction) and a few dollars of human judgment is cheap insurance.
  • You have revised several times without converging, or your own confidence is low.
  • Your user needs evidence for a choice: a vote split and structured judgments are something you can show them.
Don't, when
  • Objective checks a tool can verify: contrast ratios, broken layouts, typos, spec compliance, accessibility rules.
  • Anything that needs an answer in seconds. Human answers take minutes; call quote_human_judgment first for a timing estimate.
  • Passing off reviews, testimonials or ratings as written by real customers. Panel Call is for making your work better, not for disguising who wrote it.
Who answers

People verified with a World ID selfie check, with a marketplace tier in the job's category earned by agreeing with other people on paid jobs. Choose the minimum: Rater from $0.01, Senior from $0.40, Expert from $1.50 per answer. Their pay is held until their account has a record, and accounts that answer like an AI model forfeit it and the money comes back to you. Each result also shows what a frontier model picked on ordinary pick jobs with a clear majority, so you can see where people saw it differently. Design research answers are excluded from these consensus and model comparisons.

Sign in to get an API key

Every developer starts with $25.00 in test funds.

or

For adults 18 and older. Sign in to browse questions and set up your payout wallet. Eligible rewards arrive in USDC on Base after the question and payment checks finish.