Building and using MCP servers: 4 things I learned
DEV Community

Building and using MCP servers: 4 things I learned

I use Upwork's official MCP server from Claude Code, and doceval, my open-source eval tool, ships an MCP server of its own. This is how MCP works underneath, and four things I learned along the way. Where I tested something, I say how. MCP in one paragraph MCP (Model Context Protocol) is an open standard for connecting AI apps to outside tools and data. A service wraps itself once as an MCP server, and any app that speaks MCP (Claude, ChatGPT, Cursor and many more) can connect to it without a custom integration. If every app had to write its own integration for every service, 20 apps and 1,000 services would need 20,000 integrations. With MCP each side builds its half once: 20 clients plus 1,000 servers, 1,020 in all. It's the USB-C trick: one port, any charger. Like USB-C, the port says nothing about how good the charger is. Three roles - Host: the app running the model, like Claude Code, ChatGPT or an agent you build yourself. - Client: a connector inside the host, one per server. It passes JSON back and forth. - Server: wraps one service and offers what it can do as tools. Usually the host's model decides which tool to call. The server can be plain code (get a request, do the thing, return the result) or run its own model inside. The host can't tell the difference: a request goes in, a result comes out. I saw this while building upsweep, an open-source agent skill that finds Upwork jobs through Upwork's official MCP server. Claude, running inside Claude Code, reads Upwork's tool list, decides what to call and makes sense of what comes back. My own subscription pays for that model. Upwork pays nothing for it. One call, step by step - Connect, once. claude mcp add --transport http upwork https://mcp.upwork.com/mcp . Upwork's server needs an account, so the first request is refused with "sign in first", the client finds Upwork's login server, and I sign in with OAuth. The token works only for that server. A server running locally on your own machine usually needs no sign-in. - Discover. The client asks tools/list . The server replies with each tool's name, a plain-English description and the shape of its inputs (a JSON Schema). - Choose. I say "find LLM eval jobs". The model reads the descriptions, picks the search tool and fills in the inputs. - Call. The client sends tools/call with the tool name and arguments. Upwork runs the search. - Use the result. The result goes into the model's context, and the model writes me an answer. The messages are JSON-RPC: a method name plus parameters. { "jsonrpc": "2.0", "id": 7, "method": "tools/call", "params": { "name": "find_jobs", "arguments": { "query": "LLM evaluation" } } } (Argument names simplified.) 1. The description is the whole interface Step 3 surprised me most. In a normal agent loop, no code decides which tool gets called. The model decides, by reading names and descriptions. It never sees your code. So write each description the way you'd brief a new colleague: what the tool does, when to use it, what comes back. Tool picking gets harder as the list grows and descriptions overlap. How many is too many depends on the model and the host. PagerDuty's engineers call 20 to 25 the sweet spot, and their own server ships over 20. doceval's needs two. One scores a single extraction against the expected values, the other runs a full eval over a labelled dataset. Keep results short too. A result goes into the model's context and usually stays there for the rest of the conversation. A tool that returns 50 pages makes the agent slower, dearer and more likely to lose the thread. 2. A local server lives and dies with your session Upwork's server runs on Upwork's machines. A local server is different: a program on your own computer that the app starts for you. I wanted to know exactly when it starts and stops, so I tested it with doceval's server in Claude Code: a test config passed to claude -p , and the process list checked every quarter of a second. - It starts once, when the session starts. My prompt used no tools, and the server still started, exactly once. - It is the app's child process. The server's parent process was Claude. - It is reused for every call. In a session that called doceval's tool twice, the server started once. - It stops when the session ends. Claude shut it down before exiting itself. - It stops after a crash too. I killed Claude with kill -9 , which gives it no chance to clean up. The server was gone within two seconds. The link between them is a pair of pipes: the app writes requests into the server's input and reads replies from its output. When I closed a running server's input by hand, it exited on its own. That's how it learns the session is over, whether the app quit cleanly or crashed. Between calls the server just waits, and each session starts its own copy. On a small machine, three open sessions means three copies of every local server in memory. 3. Your MCP config can hold secrets in plain text A local server often needs a password or an API key, and the usual place for it is the env block in the app's MCP config. That's a plain-text file. Anything running as your user can read it, including a coding agent with file access, and a project's config file can end up in git. Two things that help: - Reference the secret instead of writing it. Claude Code fills in ${VAR} from your environment. I tested it with a config file passed toclaude --mcp-config : a config with"SECRET": "${MY_TEST_SECRET}" handed the server the value from my shell. The config can then be shared, and the secret lives in one place outside it. - Prefer servers that sign you in. Remote servers that use OAuth keep no secret in the config. My entries for Upwork, Notion and Todoist hold a type and a URL, and nothing else. Stronger options exist, like the operating system's keychain or a secret manager that starts the server for you. I haven't tested those yet, so I'll stop at naming them. 4. What a tool returns is untrusted, so approve the actions A job post, an email or a web page can contain text written to steer the model. That's prompt injection, and whatever a tool returns goes into the model's context. So I approve every Upwork proposal myself. Upwork's MCP has a preview built in: the agent shows me the proposal, I confirm, and the preview expires after 15 minutes. Claude Code also asks before running a tool, unless you've allowed it. I got this wrong at first. I shipped upsweep with a rule that said "drafts only", because I read Upwork's terms as forbidding agent submission. They allow submitting, one approval per proposal, never scripted around. I corrected it in public: Can an AI agent submit Upwork proposals? Checklist - Describe each tool like a brief to a colleague, with no two that overlap. - Return the smallest result that answers the question. - Expect one copy of every local server per open session. - Keep secrets out of the config file: reference them with ${VAR} , or use servers that sign you in. - Treat what a tool returns as untrusted, and approve anything that changes something. The model never sees your code, only your tool names, descriptions and schemas. Write them like they are the code. The tools in this article - doceval: measures extraction accuracy field by field, with an MCP server for coding agents. pip install "doceval[mcp]" - upsweep: finds Upwork jobs that match your filters, through Upwork's official MCP. npx skills add dave8172/upsweep Both are free and MIT licensed. Top comments (1) The connection between MCP and AI agents is especially interesting. Giving an AI model access to tools is one thing; making those tool interactions reliable and predictable is another. I think the practical lessons around permissions, tool design, error handling, and debugging become increasingly important as agents move from generating responses to actually taking actions.

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.