Jack Gordley

Jack Gordley

Software Engineer at Grafana Labs working on AI evals, observability, and incident investigation. Previously Amazon Bedrock. Go Irish 🍀

Certified LLM-free, hand-crafted, artisan writing.

Recent writing

Recent pieces from this blog and external articles I wrote for companies or projects.

/gordles.io/Blog

Hand-crafted, artisan writing

/Grafana Blog/External

How to build a trust platform for your agent with Grafana Agent Observability

/Grafana Blog/External

Introducing o11y-bench: an open benchmark for AI agents running observability workflows

/AWS Machine Learning Blog/External

Build reliable AI agents with Amazon Bedrock AgentCore Evaluations

/gordles.io/Blog

LLM Context-Friendly Test Suite Outputs (build-output-tools-mcp)

On this blog

Posts published here on AI systems, developer tools, agents, and side projects.

/gordles.io/Blog

Hand-crafted, artisan writing

4 years ago I never would have expected to be fooled even once by writing or media created by AI. Lately I have been more motivated than ever to do ...

/gordles.io/Blog

LLM Context-Friendly Test Suite Outputs (build-output-tools-mcp)

Post covering an MCP tool to run tests/builds and routes the output to a smaller, specialized LLM for summarization to avoid flooding the main context thread of Claude Code.

/gordles.io/Blog

The Most Cost-Effective LLM Coding Loadout for the Casual Developer

TLDR; Claude Pro with Claude Code + OpenRouter Zen MCP

/gordles.io/Blog

From Text to Action: How LLMs Became AI Agents

A mini-history lesson on how LLMs gained the ability to execute actions and the state of Agents in 2025

/gordles.io/Blog

Calvin - An Open-Source Google Calendar Assistant

In-depth project rundown of a Langchain assistant that can manage your Google Calendar

External writing

Articles I wrote or co-wrote for a company, product, or project outside this blog.

/Grafana Blog/External

How to build a trust platform for your agent with Grafana Agent Observability

A practical guide to using agent observability to understand behavior and build trust in AI agents.

/Grafana Blog/External

Introducing o11y-bench: an open benchmark for AI agents running observability workflows

Open benchmark for evaluating AI agents on observability workflows.

/AWS Machine Learning Blog/External

Build reliable AI agents with Amazon Bedrock AgentCore Evaluations

A practical walkthrough of agent evaluation across development and production.

Game dev projects

A few game jams I've participated in.