Tech Stack
Description
Two kinds of work kept costing more time than the code itself. The first was the paperwork around a change: commit messages, PR descriptions, and QA instructions written by hand, differently each time, so reviewers and testers had to reconstruct intent from the diff. The second was production triage — an error appears in Datadog and someone spends hours working out which service, which path, and which change is responsible before a fix can even start.
For the first, I built Claude Code skills that read the current session and the Jira ticket, pulled in over MCP, and produce a commit message, a PR description, and step-by-step QA instructions in one consistent shape. Because the template is the same everywhere, the skills are reusable across frontend and backend, client and service repositories, and a reviewer knows where to look in any PR.
For the second, I connected Claude Code to Datadog through the Datadog MCP server, so the agent pulls an error's logs, traces, and metrics itself instead of waiting for me to copy them in. On top of that I built a diagnosis skill: it gathers the telemetry, investigates the issue in the codebase, identifies the likely root cause, and recommends fixes, which I review before applying. Diagnosis that typically took two to three hours now takes fifteen to thirty minutes, and the fix is ready to go into a pull request in the same session.
The two halves feed each other: the diagnosis skill produces the fix, and the authoring skills turn it into a standard PR and a QA plan. The skills propose and I decide — every recommendation is reviewed, tested, and owned by me before it ships.
For changes that span the frontend and the backend, I go a step further and run them as a multi-agent workflow. I write one markdown spec per change: the context, the Jira ticket (read through the Atlassian MCP server), the feature or bug, its requirements and constraints, and the testing instructions. The spec then assigns work by repository — one agent implements the backend, another the frontend — and a separate tester agent verifies that both sides meet the requirements, runs the tests, and starts the application so I can check the change myself.
The tester is always its own agent, deliberately. An agent that wrote the code is a poor judge of whether it works, so verification is kept independent of implementation, the same way code review is kept separate from authorship. The tester's report is a claim, not a verdict: I confirm it against the running application before deciding to ship.
- Designed a multi-agent workflow for cross-repository changes: one spec (context, Jira ticket via Atlassian MCP, requirements, constraints, test plan) drives separate frontend, backend, and tester agents.
- Kept verification independent of implementation with a dedicated tester/verifier agent that checks requirements, runs tests, and launches the app for manual review.
- Treated every agent's output as a claim to verify: I test the running application myself and make the final ship decision.
- Built reusable Claude Code skills that generate commit messages, PR descriptions, and QA instructions from session context and the Jira ticket via MCP.
- Standardized PR and QA handoff across frontend and backend repositories with one consistent template.
- Connected Claude Code to Datadog through the Datadog MCP server and built a diagnosis skill on top: the agent pulls the error's logs and traces itself, then returns root-cause analysis and recommended fixes, reviewed by me before anything ships.
- Cut production error diagnosis from 2–3 hours to 15–30 minutes, with the fix ready for a pull request in the same session.
- Designed the skills to be repo-agnostic so they carry over to new applications without rework.
- Kept a human in the loop: the skills propose, and I validate, test, and own every change.