No. The assistant you already talk to can read and act inside much of the software you already pay for, without that software shipping a single AI feature of its own, and the catalog on this site is the filter for which software that is. Four open pieces made it true: the MCP protocol, now at spec revision 2026-07-28 (https://modelcontextprotocol.io/specification/latest, checked 2026-08-22); the Agent Skills file format, which 46 different products now read (https://agentskills.io/clients, counted 2026-08-22); assistant plugins; and models that can hold a longer task. The work in front of you is picking one task small enough to check and pointing your assistant at it, which is a different job from watching a vendor roadmap.
Where to start, by what you use: the tools index carries the per-application pages, and Notion, Slack, HubSpot and Asana are the four with the most traffic behind them. If you want the shortest version of which applications have a server published by their own vendor, that is the best MCP servers page.
Why does the software I already pay for feel like more work than it saves?
Because most of what you do inside an app is not the work, it is the operating of the app. Opening the right view, remembering which of two spreadsheets is the live one, copying a number out of one system and typing it into another, closing the dialog that appears every time. People hate using their computers, and they are not being unreasonable. The hatred is aimed at the operating, not at the work.
There is a second reason, and it is less obvious. The interface you see is a curated slice of what the app can actually do. In a snapshot of Composio’s connector catalog taken 2026-08-14 and queried 2026-08-22, the median application this site publishes exposes 19 distinct actions, and the deepest, Canvas, exposes 574. Zendesk exposes 452. Almost nobody who uses Zendesk every day could name fifty things it can be asked to do, because the menus do not offer them. The capability was always there. The path to it went through a developer.
Those are actions retrieved from the connector catalog, which is the count this site’s own catalog file publishes. The catalog’s own listing advertises a slightly lower figure for 111 of the 1,000 applications, 565 actions lower in total, so a number one or two off the ones here is the other column rather than a contradiction.
What changed, and when did it change?
Four things, the earliest in November 2024, and none of them required an app vendor to build AI into its own product.
A protocol for tools, under neutral stewardship. Anthropic published the Model Context Protocol on 2024-11-25 as an open standard for connecting AI applications to outside data and tools (https://www.anthropic.com/news/model-context-protocol). In December 2025 it donated the protocol to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded with Block and OpenAI, where the existing maintainers keep full autonomy over technical direction (https://blog.modelcontextprotocol.io/posts/2025-12-09-mcp-joins-agentic-ai-foundation/, retrieved 2026-08-22). That is the part that makes this worth learning rather than watching. A feature belongs to one vendor and can be repriced, moved behind a higher plan, or retired. This one is not any single company’s to withdraw.
The specification is now on its 2026-07-28 revision, maintained in public (https://blog.modelcontextprotocol.io/posts/2026-07-28/, retrieved 2026-08-22). It also states a security principle worth reading exactly: hosts must obtain explicit user consent before invoking any tool. The same section says that MCP cannot enforce these principles at the protocol level and that implementors SHOULD build consent and authorization flows into their applications (https://modelcontextprotocol.io/specification/latest, checked 2026-08-22). So consent is written into the standard as a principle and delivered by whoever builds the client. Both halves matter.
A file format for procedures. A skill is a directory holding a SKILL.md
file with a name, a description, and instructions, plus optional scripts and
reference material. The format was originally developed by Anthropic and released
as an open standard, published at agentskills.io, which is open to outside
contributions (https://agentskills.io/home, checked 2026-08-22). A count of the
client showcase on 2026-08-22 found 46 named products reading it, including
Claude, ChatGPT and Codex, Gemini CLI, GitHub Copilot, VS Code, Cursor, JetBrains
Junie, and Kiro (https://agentskills.io/clients, counted 2026-08-22). OpenAI’s
own documentation for Codex says its skills “build on the open agent skills
standard” and points at that same specification
(https://learn.chatgpt.com/docs/build-skills, checked 2026-08-22). This is the
part most people do not know: the interoperability question is settled at the
file level, in the vendors’ own docs.
Plugins, with a caveat. Claude, ChatGPT, and Microsoft Copilot Studio all
take plugins or their equivalent today, and Microsoft made MCP integration
generally available in Copilot Studio in May 2025
(https://www.microsoft.com/en-us/microsoft-copilot/blog/copilot-studio/model-context-protocol-mcp-is-now-generally-available-in-microsoft-copilot-studio/,
published May 2025; the exact day was not confirmed). Of the four routes into an
app, the plugin route looks like the weaker one. In the AI answer channel the
term chatgpt plugin fell from 755 to 216 over twelve months, down 71 percent,
while every MCP term rose (DataForSEO AI Keyword Data, US English, measured
2026-08-21; see the caveat at the end of this page). Reported here because it is
what the measurement says, not because a standard has been retired, and search
volume is a measure of curiosity rather than of whether a route works.
Models that can hold a longer task. METR measures how long a task an AI agent can finish, and in Time Horizon 1.1, published 2026-01-29, the estimate for the strongest model in the set was 320 minutes with a confidence interval of 170 to 729, on a doubling time of 196 days over the full period and about 131 days counting only from 2023 (https://metr.org/blog/2026-1-29-time-horizon-1-1/). Read the qualifier before you read the number: that is the length at a 50 percent success rate. Half. Nothing you care about should be handed over at half.
Why won’t the app vendors just build this themselves?
Their incentive points the other way, and the evidence for that is in their pricing pages. Of the 589 applications this site publishes, 175 have some in-product AI, 95 have something limited, 181 have none, and 138 have not been checked yet (catalog snapshot, queried 2026-08-22). Notion states that Notion AI runs on Business and Enterprise plans, with Free and Plus getting a capped number of complimentary responses (https://www.notion.com/help/notion-ai-faqs, verified 2026-08-22). That is the pattern rather than an exception. The feature arrives, and it arrives on a plan you are not on.
Making a product reachable from outside is a separate, slower decision, and more
vendors have made it than the discussion suggests. The protocol’s public registry
is a paginated open API whose limit parameter caps at 100 results. Walking all
243 pages with version=latest on 2026-08-22 returned 24,221 distinct server
names (https://registry.modelcontextprotocol.io/v0/servers, full pagination
2026-08-22). Counts near 7,100 that circulate, including earlier ones published on
this site, came from a walk that stopped before the end.
Joining that full registry against the 1,000-application connector catalog by name on 2026-08-22 matches between 294 and 318 applications depending on how strictly names are compared, and between 165 and 180 of the 589 this site publishes. Those are name matches rather than verified vendor servers. The verified per-application check is the census, and it is not finished. What has been checked one application at a time is on the best MCP servers page: of 41 widely used business applications queried individually on 2026-08-22, 20 had a server published by the vendor itself and 21 did not.
So the reachable surface is wider than the AI-feature surface, and the honest answer for any single application is still to look it up rather than to reason about it.
What can my AI actually do inside an app I already pay for?
More reading than changing, and the reading is where most of the surface lives. Across the 589 published applications there are 26,826 distinct actions. Counting an action as a read when its name, with the application prefix removed, contains get, list, search, find, retrieve, fetch, or read as a separate word, 13,303 of them qualify, about 50 percent (catalog snapshot, queried 2026-08-22). A different word-boundary rule moves the count by a few hundred and does not move the shape.
The other number worth knowing is smaller than people expect. Only 34 of those 589 applications can start work without being asked, and only 38 across all 1,000 entries in the snapshot. Almost nothing in an ordinary stack acts on its own. In practice, “agentic” means your assistant does what you asked, inside the app, at the moment you asked it, and then stops.
That is a less exciting sentence than the category usually gets, and it is the one that makes the next section possible.
How much of a job should I hand to my AI?
That depends on the job, not on how much you have handed over before. There are four shapes this work takes. They are not four levels, there is no order to them, and nobody graduates. People who do this well keep using all four permanently, because each one is wrong for exactly what the others are right for. This is the canonical version of the table; the other guides on this site carry only the row that matters to their own subject.
| Shape | What it looks like | Right for | Wrong for |
|---|---|---|---|
| Reading | Ask your AI to go look at something in a tool and tell you what it found | Finding out what is actually in there, before you decide anything | Anything that changes a record. A read cannot be the step that commits |
| One bounded change | Ask it to do a single thing in a single tool | Work you can check by looking at the result, and undo if it is wrong | Work whose correctness you cannot see. If you cannot check it, you have moved the error, not removed it |
| One handoff | Connect two tools for one transfer you already do by hand | A transfer whose rules you can state out loud in two sentences | Judgment that depends on context only you hold. The handoff will apply the rule and miss the exception |
| One owned procedure | Write the procedure down in a form your AI can execute, keep it in a file you control | Work you repeat often and can specify exactly | Work whose rules are still changing. You will spend more rewriting the file than you saved |
All four are the same act at different scopes: saying exactly what you want, in enough detail that someone who was not in the room could do it. Most people have never had to write that down, which is not a personal failing. Writing a task down clearly enough to hand off is management work, and most people were never taught it, because that training goes to the few identified early as worth training. The specification is the skill, and it is a slower thing to build than a tool is to install.
Which is where METR’s 50 percent belongs. The capability curve is real and measured. It is also measured at a success rate nobody would accept from a colleague, on 228 tasks drawn from RE-Bench, HCAST and shorter novel software tasks, primarily software engineering, machine learning, and cybersecurity work (https://metr.org/blog/2026-1-29-time-horizon-1-1/, published 2026-01-29). A rising curve tells you the ceiling is moving. It does not tell you the thing in front of you today is checkable, and checkable is the only test that matters.
Do I have to be good at the app before I can use AI on it?
No, and the belief that you do is the most expensive wrong idea in this subject. People skip AI in Excel because they are not Excel experts, which has the dependency backwards. What the assistant lacks is not skill with the software. It is your knowledge of your own work.
What it needs from you is answers, not expertise. Four questions cover most of it, and they are the same four used on every guide on this site.
- What is this task’s one job?
- What does a good result look like, specifically enough that someone else could tell?
- What must never happen without my approval?
- What supervision does this task need?
None of those is a technical question. All of them are things you already know and have never had to say out loud.
When do I actually do this?
Not inside the work, which is the reason most people who intend to never start. The reading, the one bounded change, and the writing down all happen in the same place the work happens, competing with it, and the work always wins. Nobody sets aside time to stop working in their work and work on it, and being told that AI helps without being shown where makes that worse rather than better.
The unit that fits is one task and one sitting. Pick something you did last week that you will do again next week, spend the time you would have spent doing it once on getting your assistant to read the inputs instead, and stop there. That is a complete outcome. It is also the only version of this that survives a normal week.
Do I have to be technical to set this up myself?
Less than you would have needed two years ago, and the reason is that the pieces became files instead of code. A skill is a folder with a Markdown file in it, and the Markdown file needs a name, a description, and instructions (https://agentskills.io/home, checked 2026-08-22). If you can write a document, you can write one. Keeping it in a Git repository, which used to be the part that scared people off, is now closer to using a shared drive that remembers every version. For the longer version of that, see what Claude skills are and how to write an AI-first SOP.
The connection is the part that used to be genuinely developer work, and it is where a managed connector differs from the old approach of pasting an API key into a configuration file. A managed connector holds the credential on its own side, refreshes an OAuth token before it expires, and lets a connection be switched off in one place without editing any file, while the provider’s own security settings remain able to revoke it outright (https://docs.composio.dev/docs/tools-direct/authenticating-tools, checked 2026-08-22). A key pasted into a config file does none of that, and it keeps working, silently, until somebody remembers it is in there. The connector is plumbing. It is worth understanding for exactly one reason, which is that plumbing you do not understand is plumbing you cannot turn off.
Is a skill the same thing as an SOP?
Close enough that treating them as the same thing is the useful move. An SOP written for a person leaves out everything the person already knows: which system is authoritative, what the exception looks like, who to ask. An SOP written for an agent has to put all of that back in. Doing that translation is how most people discover that their procedure was never actually written down, only performed.
The two standards are converging on purpose. The MCP project runs a Skills Over MCP working group whose stated long-term goal is interoperable skill distribution across MCP servers and clients, co-led from Anthropic and Nordstrom with participants from Google, GitHub, AWS, Databricks, Bloomberg, and Saxo Bank, meeting weekly. Its charter was formalized 2026-04-14 and it was converted from an interest group to a working group on 2026-04-16 (https://modelcontextprotocol.io/community/working-groups/skills-over-mcp, checked 2026-08-22).
Almost nobody is searching for this, on either instrument, and the figures are in the sources at the foot of this page. The section is here because it is where the work ends up, not because a keyword asked for it. The two guides that carry it in full are what Claude skills are and how to write an AI-first SOP.
Whose job is this, mine or IT’s?
Both, on different questions, and the line is cleaner than it looks. You can do on your own: reading anything you already have access to, one bounded change in a tool where the account is yours, and usually a handoff between two tools you personally log into. You cannot do on your own: anything needing an administrator to approve a connection at the workspace level, anything writing to a system of record other people depend on, and anything where the honest answer to “who else does this affect” is more than you.
Finding out which one you are looking at is a five-minute question, and it is worth asking before you build. When you do ask, ask for something specific. “Can I connect my assistant to my own Notion workspace, read-only” is a question an administrator can answer today. “Can we do AI” is a question that goes into a queue.
There is a version of this that is genuinely a group problem. Every department’s leadership tends to treat AI adoption as another department’s project, which is how a company ends up with nobody owning it and everybody waiting. Knowing which side of your own line a task falls on is the part you can settle this week.
This page is about what one person can do with the software they already have. If the honest answer to your situation is that it needs a team, a budget, and a decision about who owns the output, that is different work with different failure modes, and nothing here should be read as a claim about it.
When is waiting actually the right call?
When the thing you would be waiting for is a control you cannot build yourself. Four cases, and they come up often enough to name.
The app is not reachable. This site’s catalog is a filter, not a directory of all software. Thousands of products have no API worth integrating, no managed connector, and no MCP server, which is why they never appear here at all. If yours is one of them, there is nothing to point an assistant at, and the honest options are to wait or to change tools.
The output is regulated or audited and nobody has decided who owns it. The MCP specification puts consent at the individual tool call, which protects the moment of action. It says nothing about how the result gets attributed six months later in a review. That decision is organizational, and doing the work first does not make it.
The rules are still moving. A written procedure that gets rewritten monthly costs more in maintenance than it returns. Do it by hand until the rule holds still.
You cannot check the result. This is the one that disqualifies the most tempting tasks. If you have no way to look at the output and know whether it is right, handing the task over relocates the error into a place you will find it later, at a worse time.
What could a careful check not establish?
Ten things, and the largest is whether the one hard measurement on this page describes the kind of work the people reading it actually do. None of the ten is a reason to wait.
Whether any of this saves time, or how much. No measurement was run. Any time saving attached to a connected assistant, including any figure this site produces later, is an estimate until it is measured against the specific task.
Whether the capability curve applies to office work. METR’s time horizon is measured over 228 tasks drawn from RE-Bench, HCAST, and shorter novel software tasks, primarily software engineering, machine learning, and cybersecurity (https://metr.org/blog/2026-1-29-time-horizon-1-1/, published 2026-01-29). Of the 31 tasks estimated at eight hours or longer, only 5 have measured human baseline times; the rest use estimates. No equivalent published curve was found for month-end close, contract review, or scheduling.
How many of the 589 a person can turn on unaided. The 589 published applications were not checked for how many are reachable without an administrator enabling something first. That count does not exist yet, and it is the gap between this page’s figure and what a reader can do this afternoon.
How many reachable apps there actually are. The catalog snapshot stops at exactly 1,000 because that is the ceiling of the tool that produced it, and the snapshot’s own metadata records the truncation in writing. So 1,000 is a floor and the real figure is unknown.
Whether an action changes anything. The read proportion classifies 26,826 actions by the shape of their names, not by reading what each one does. Some calls that look like reads have side effects, and some that look like writes are idempotent. Treat it as a rough shape, not an audit.
Which action count is the right one. Two columns in the catalog disagree for 111 of the 1,000 applications, 565 actions in total. This page uses the retrieved count, which is what the site’s own catalog file publishes. Nobody has decided which column is canonical.
What “official” means in the MCP registry. Two things settle it in practice: a
reverse-DNS namespace on the vendor’s own domain, or an io.github.<org>
namespace where that organization owns the product’s canonical repository. Neither
is verification of quality, maintenance, or safety, and a third party can publish
a server for a product it has nothing to do with.
Whether the plugin route survives. The 71 percent twelve-month decline in
chatgpt plugin is measured in a channel that DataForSEO derives from People Also
Ask data. It is a modeled rate, not observed assistant traffic. A proxy pointing
down is weaker evidence than it looks, and it measures curiosity rather than
whether the route works.
Whether a bundle travels. A single SKILL.md file is read by 46 products
today, and Anthropic’s own help center says the format is not locked to Claude. A
bundle of a skill plus a server plus its configuration is a separate question, and
the Skills Over MCP working group puts plugin and bundle packaging explicitly out
of its own scope
(https://modelcontextprotocol.io/community/working-groups/skills-over-mcp,
checked 2026-08-22). No cross-client execution test was run here.
Whether any of the vendor verdicts are still true. All 83 vendor-documentation checks behind this site’s AI verdicts carry a single verification timestamp, 2026-08-22. Not one of them has been re-checked since. A dated claim is more honest than an undated one and it is not the same as a current one.
What this page checked, and when
| Claim | Source | Date |
|---|---|---|
| MCP published as an open standard | https://www.anthropic.com/news/model-context-protocol | 2024-11-25 |
| Donation to the Agentic AI Foundation under the Linux Foundation; maintainers keep technical autonomy | https://blog.modelcontextprotocol.io/posts/2025-12-09-mcp-joins-agentic-ai-foundation/ | published 2025-12-09, retrieved 2026-08-22 |
| Specification revision 2026-07-28 published | https://blog.modelcontextprotocol.io/posts/2026-07-28/ | published 2026-07-28, retrieved 2026-08-22 |
| Consent as a security principle; MCP cannot enforce it at the protocol level; implementors SHOULD build consent flows | https://modelcontextprotocol.io/specification/latest | checked 2026-08-22 |
Agent Skills format developed by Anthropic, released as an open standard, open to contributions; a skill is a directory with a SKILL.md file | https://agentskills.io/home | checked 2026-08-22 |
| 46 named products read the Agent Skills format | https://agentskills.io/clients | counted 2026-08-22 |
| OpenAI Codex skills build on the open Agent Skills standard | https://learn.chatgpt.com/docs/build-skills | checked 2026-08-22 |
| Skills Over MCP working group, charter and scope | https://modelcontextprotocol.io/community/working-groups/skills-over-mcp | checked 2026-08-22 |
| MCP generally available in Microsoft Copilot Studio | https://www.microsoft.com/en-us/microsoft-copilot/blog/copilot-studio/model-context-protocol-mcp-is-now-generally-available-in-microsoft-copilot-studio/ | published May 2025, exact day not confirmed |
| 320 minute time horizon, CI 170 to 729, 196 day doubling, 228 tasks, 31 long tasks with 5 human baselines, 50 percent success qualifier | https://metr.org/blog/2026-1-29-time-horizon-1-1/ | published 2026-01-29 |
| Notion AI on Business and Enterprise plans | https://www.notion.com/help/notion-ai-faqs | verified 2026-08-22 |
| Managed connector credential handling and revocation | https://docs.composio.dev/docs/tools-direct/authenticating-tools | checked 2026-08-22 |
Public MCP server registry, 24,221 distinct server names, limit caps at 100, 243 pages at version=latest | https://registry.modelcontextprotocol.io/v0/servers | full pagination, 2026-08-22 |
Application counts, action counts, in-product AI verdicts, and trigger counts on this page come from a snapshot of Composio’s connector catalog, built 2026-08-14 and queried 2026-08-22. The registry-to-catalog name join was run against that snapshot on 2026-08-22.
Google search volumes and AI-channel figures come from DataForSEO, United States,
English, measured 2026-08-21 and 2026-08-22. chatgpt plugin fell from 755 to 216
in the AI channel over twelve months. ai first sop returns no measurable Google
volume and no measurable AI-channel volume; agentic sop returns 6 in the AI
channel. The AI-channel numbers carry one
caveat that travels with every use of them. DataForSEO derives
ai_search_volume from People Also Ask data in the search results page. It is a
modeled rate, not a count of assistant queries, and it is not summable: nineteen
phrasings of the same Notion question all return the same figure. Read it as a
second directional instrument beside Google volume, never as a measurement of what
assistants were actually asked, and never as evidence about whether a route
works.