AWS ships 38 healthcare agent skills with a benchmark, and Bedrock's own record lights none

AWS ships 38 healthcare agent skills with a benchmark, and Bedrock's own record lights none

AWS has published the first agent-skills release I have seen that arrives with a measured result instead of a demo. In Improving HCLS AI reasoning with open-source agent skills, Michael Hsieh opens with the failure: “AI agents built on foundation models often misapply healthcare and life sciences decision frameworks, even when they’ve seen the guidelines in training and in the system prompt.” The agent “will cite the correct framework but misapply evidence categories, skip population frequency thresholds, or hallucinate computational predictor scores.” The answer is 38 open-source skills across 11 healthcare and life sciences domains, genomics, drug discovery, claims and risk adjustment, medical imaging, trial design, released under MIT-0 in the awslabs repository, each a markdown file with YAML frontmatter, decision frameworks, parameter tables, and validation criteria.

The numbers are AWS’s own, on its own evaluation, and they are reported with effect sizes, which is more than most vendor benchmarks manage. Across 410 prompts on each of two agent configurations, the Kiro CLI and a Strands agent, the skilled agent won 69.5% to 85.9% of head-to-head comparisons, with the strongest gain on critical thinking at an effect size of 0.65 to 1.03. Skills cut the standard deviation of response scores by 52% to 62% within a domain, and helped most where the base agent struggled most. What survives the vendor framing is the design argument: “Skills are auditable, portable, and straightforward to maintain. Every decision criterion is human-readable in markdown format, not hidden in model weights.” The cost is stated too. Loading all 38 into one agent context consumes about 80,000 tokens, which is why the post routes them through a multi-agent architecture instead.

The catalog has to make a distinction the post does not. These are skills for the domain, not for the API: they teach an agent how to interpret a variant or sequence an imaging pipeline, not how to call Bedrock. The Amazon Bedrock provider page lists 8 API pages, and the ones the deployment path touches are the Amazon Bedrock Agent API and the Amazon Bedrock Agent Runtime API. AgentCore, the runtime the post says to move to for production, still has no API page under the record, the same gap this series found on the Bedrock MCP Apps story six days ago. The agentic access profile maps 12 operations, 6 of them acting.

The Kin Score is 55.0, strong band, carried by access clarity at 79.5 and discoverability at 66.1, with contract quality at 57.8, developer ergonomics at 46.4, and contract governance at 23.5. The Agent Readiness score is 23.0, agent-aware. The agent skills dimension is unlit, and that is the finding: AWS has just shipped 38 benchmarked skills for agents running on Bedrock, and the catalog can find no skill that teaches an agent to use Bedrock itself. Auth clarity, idempotency, the MCP server, and every identity dimension are unlit too. AWS has written the best evidence yet that a markdown file can change what an agent does. The record is waiting for the file that would change what an agent can do with the API.

← The Atlassian estate counts ten members and 27 APIs, and leaves out Trello, Loom and 183 APIs on its own parent profile
Dynatrace scores 85.7 and ships its own MCP server, but 14 of the 23 contracts we hold never say where to call →