Step 1: write the task down
Write one task in full. Note the goal, the inputs you have, the output you need, and how often it should run. A vague goal makes every builder look good. A written task lets you test each one the same way.
Step 2: check the tools
List the outside tools the task needs. Then look for each one on the vendor’s own integrations page. Count only what the vendor names. A tool that appears only in a marketing sentence is not a tool you can rely on.
Step 3: check how the setup works
Find out whether you can start from a plain-language description, whether there is a visual builder, whether there are templates for your kind of task, and whether outside AI agents such as Claude Code or Cursor can connect. Then build a small version of the task and see how long the setup takes.
Step 4: check oversight
Ask four questions. Can the agent run on a schedule, or keep working when you close the app? Does it ask a person before an outside action, such as sending an email or changing a record? Can you open a run history? Can you see which tool each step used, and why? If your team will share the agent, check the team plan too.
Step 5: check the billing
Look for a free way to start with no card. Look for prices on a public page. Check whether one balance covers all the tools, or whether you pay each tool provider separately. Find out whether you can see what each run used. Ask the vendor what one run of your task costs. If the answer is not on the page, treat that as a gap.
Step 6: run one task and check the result
Run the written task. Compare the output with the goal you wrote in step one. Check each source the agent used. A builder that scores well on documentation can still produce weak results for your work, so the test matters more than the ranking.
A short comparison sheet
| Question | What a good answer looks like | Red flag |
|---|---|---|
| Which tools can it reach? | A named list on a public page, with the categories you need | Tool names only in a sales call |
| How do I start? | A plain-language start, templates, or a visual builder | Setup needs a call before any test |
| Does it wait for approval? | Named approval steps for outside actions | No mention of approvals |
| Can I see what it did? | A run history and a step-by-step view | Only a final answer, no trace |
| What does it cost? | A free start with no card, and published plans | Prices shown only after sign-up |
What the awards cover, and what they do not
The awards score what each vendor publishes about the six steps above. They do not measure how well the agent does your work. Read the methodology to see the checks, and the Institute standards for the full list of what a vendor should publish.
Frequently asked questions
Should I pick the top-ranked builder?
Only if its checks match your task. A ranking measures what a vendor publishes. It does not measure how well the builder does your work, so run your own test before you commit.
What is the first thing to test?
Run one small task that you already know the answer to. If the result is wrong, or the sources are missing, that tells you more than any score.
Do free plans change?
Yes. Check the vendor’s pricing page on the day you sign up. The awards record plan details on the review date only.
Published 2026‑10‑11 · Panel review of public documentation, October 2026