The throwaway-account rule
- Every agent is installed on an account iBrothers owns.
- Seeded with dummy data, and nothing else.
- Never a customer's account, never a venture's.
- Access revoked afterwards; the account wiped before the next review.
Where a vendor's claim meets a real install: the same twelve checks on every agent.
Most directories repeat what a vendor says. We install every agent on a throwaway account we own, write down what it actually asks for beside what its documentation says it needs, and publish the method for every check so you can judge it yourself.
From install to profile, each with the method a reviewer follows, and what a reader is told when it passes and when it does not. The ones marked Observed happen on the throwaway account.
Look the vendor up in the company registry for its stated jurisdiction. The registered name must match the vendor named on the profile, or the profile must say why.
Vendor: :name, :jurisdiction
No matching company found in the registry
Check WHOIS for the registration date, load the site over HTTPS, and confirm the footer or terms name the same entity as the registry lookup.
Domain since :year; site live
Site did not load, or names a different entity
Install the agent or reach its trial without talking to sales. A waiting list is recorded as such rather than failed, but it is not eligible to be listed.
:state
No installable product or trial found
The vendor's own documentation must list the OAuth scopes, API keys, roles or file paths the agent needs. Each one goes on the permission worksheet as declared.
Declared: :count permissions
The documentation does not list the permissions it needs
Install the agent on a throwaway account owned by iBrothers, seeded with dummy data. Record every scope the consent screen actually asks for and compare it, scope by scope, with the declared list. A scope requested but not declared cannot pass.
Requested :count; :match
Could not be observed, or requested permissions it did not declare
The vendor must state where data goes, how long it is kept, whether it is used for training, and which sub-processors see it. We check the statement exists, not that it is true.
Data-flow statement present
No statement of where data goes or how long it is kept
On the throwaway account, give the agent a task and watch: does it act without confirmation, or is there an approval step before anything is sent, written or paid?
:observed
Autonomy could not be observed
Revoke the agent's access, stop it, and request deletion of its data. Count the steps, and confirm the access is actually gone from the provider's side afterwards.
Revocable in :steps steps; verified
Access could not be revoked cleanly
The profile must state which models are used, from which provider, whether they are the vendor's own or an API, and how the agent is evaluated. We check the questions are answered, not that the answers are true.
Provenance stated (content not verified)
Models, providers or evaluation not stated
A privacy policy and terms must be present, dated, and name the entity. A claimed certification must link to the issuer or a letter, not to a logo.
:summary
Policies missing, undated, or name no entity
A pricing page must be reachable without contacting sales. "Contact us" is recorded as not public rather than failed.
Pricing public
Pricing not public
There must be a named support channel, and a documented process for incidents and for rolling back a bad release.
Support channel and incident process stated
No support channel or incident process found
Never checked: accuracy, "best", return on investment, "hallucination-free". Those stay the vendor's own claims, shown with their evidence or marked Unverified.
Every permission, written the same way. Every scope an agent asks for is described in the same words on every profile, with how much it lets an agent do. These are the ones we have wording for so far.
20 scopes
12 scopes
6 scopes
https://mail.google.com/
High risk
Full access to the mailbox, including permanent deletion
Almost no agent needs this; ask why.
calendar.events
Medium risk
Can create, change and delete events on your calendars
calendar.readonly
Medium risk
Can see every event on your calendars, with its details
contacts.readonly
Medium risk
Can read your contacts
drive
High risk
Can read, change and delete every file in your Drive
drive.file
Low risk
Can see and edit only the files it creates or you open with it
drive.readonly
High risk
Can read every file in your Drive
gmail.compose
High risk
Can write drafts and send email as you
gmail.modify
High risk
Can read, label, archive and bin email
gmail.readonly
High risk
Can read every email and attachment in the mailbox
gmail.send
High risk
Can send email as you
Calendars.Read
Medium risk
Can see every event on your calendars
Calendars.ReadWrite
Medium risk
Can create, change and delete calendar events
Files.Read.All
High risk
Can read every file you can open in OneDrive and SharePoint
Files.ReadWrite.All
High risk
Can read, change and delete every file you can open in OneDrive and SharePoint
Mail.Read
High risk
Can read every email in the mailbox
Mail.ReadWrite
High risk
Can read, change and delete email
Mail.Send
High risk
Can send email as you
offline_access
Medium risk
Keeps its access while you are not using it
User.Read
Low risk
Can read your basic profile: name, email address and photo
AdministratorAccess
High risk
Can do anything in the account, including deleting it
No agent should need this; treat it as a refusal to scope.
ce:GetCostAndUsage
Low risk
Can read the account's cost and usage reports
iam:PassRole
High risk
Can hand an IAM role to a service - a common route to wider access
ReadOnlyAccess
High risk
Can read almost every resource and setting in the account
s3:GetObject
Medium risk
Can read objects in the S3 buckets it is given
s3:PutObject
Medium risk
Can write objects to the S3 buckets it is given
admin:repo_hook
High risk
Can add and remove webhooks on your repositories
public_repo
Medium risk
Can read and write your public repositories
read:org
Low risk
Can see which organisations and teams you belong to
read:user
Low risk
Can read your GitHub profile
repo
High risk
Can read and write every repository you can reach, private ones included
A GitHub App limited to chosen repositories asks for far less.
workflow
High risk
Can change GitHub Actions workflows, and so what runs in CI
channels:history
Medium risk
Can read messages in the public channels it is added to
chat:write
Medium risk
Can post messages
files:read
Medium risk
Can read files shared in the channels it can see
groups:history
High risk
Can read messages in the private channels it is added to
im:history
High risk
Can read direct messages it is part of
users:read
Low risk
Can see the people in the workspace
Put it through the same twelve. Free, now and later. You see the profile before anyone else does.
A monthly note listing what was added and re-checked. Nothing else.
For the newsletter only, nothing else. How we handle it