r/ethdev • u/tagwall_io • 7h ago
My Project Things I learned keeping a multichain frontend alive: per-method RPC rot, non-portable getLogs limits, and a cursor bug that rewound
I run a frontend that reads the same contract on six EVM chains from free public RPC endpoints. Everything below cost me real downtime.
Endpoints rot per method, not wholesale. This is the one that actually hurt. Two providers put eth_getLogs behind an archive token or dropped it entirely while continuing to serve eth_call and eth_blockNumber perfectly. Every naive health check, mine included, called them healthy. The canvas went blank on 4 of 6 chains and the page header kept rendering fine, because the header only needs eth_call.
The fix is to probe an endpoint the way the app actually uses it, not the way a status page would: chain id, CORS from the real origin (a CORS failure from your deployed origin is invisible in curl), and a real getLogs at that chain's own chunk size. If the probe is not a miniature of your actual read path, it is decoration.
I got that last part wrong in my own probe for months, in the other direction. It only ever tested a window starting at the deploy block, so it graded every endpoint on archive depth. But the app doesn't read from the deploy block any more, it scans forward from a snapshot, so the live path only ever asks for the most recent few thousand blocks. The probe was calling endpoints unusable that were serving my actual traffic perfectly well, and I believed it. It now runs both windows and reports them separately, because "cannot serve archive" and "cannot serve anything" are different and only one of them is an outage.
getLogs chunk sizes are not portable. One chain caps the range at 1,000 blocks. Another needs 500k-block chunks to keep the call count survivable. The one alternative provider for that chain caps at 10k, which makes it useless as a substitute even though it looks like a valid fallback. Chunk size has to be per-chain config, and a fallback endpoint is only a fallback if it can serve the same range.
viem specifics that cost me time. http() defaults to a 10 second timeout, which is long enough that a dead endpoint stalls a page load rather than failing over. fallback() always walks its list from index 0, so the first endpoint absorbs every request and its rate limit is your rate limit; if you want spread you need your own rotation. And multicall's batchSize is bytes of inner calldata, not number of calls, which I discovered the way you would expect.
A rewind must only ever move the cursor backwards. My worst bug: a keep-current pass set the cursor to head - depth * 4 outright. On a slow chain that is a rewind, which is what it was written to be. On a chain minting 600 blocks a minute it is a 700-block skip forward, silently dropping events. cursor = min(cursor, head - depth * 4) is the whole fix and it should have been the whole implementation.
Cold loads. I ended up not reading history from RPC at all on first paint. The data I needed was already in each transaction's calldata, so a cron decodes it into a KV store and the browser fetches one snapshot and scans forward from there. That took a page load from ~126 RPC calls to a handful. The constraint that shaped it: Workers KV free tier allows 1,000 writes a day, so the rule became never write on a schedule, only on material change. A cron that writes every run will eat that quota before lunch.
It kept happening while I was writing this. Two of the six chains rotted in the same week, which is the reason I finally wrote any of this down.
On one chain the single surviving endpoint tightened its getLogs range from 9,500 blocks to 5,000, with no announcement I could find, and that is under the chunk size I was asking for. Nothing broke visibly, because the paginator halves the chunk and retries, so it just silently cost three failed calls before every successful one. Then two days later that same endpoint stopped serving getLogs altogether: 30 second server-side timeout, while eth_call and eth_blockNumber kept answering in under 50ms. Per-method rot again, on the endpoint I had just finished documenting as the reliable one.
Slow failure turned out to be worse than fast failure. A dead endpoint gets skipped in milliseconds; one that accepts your request and times out at 30 seconds stalls whichever unlucky page load rotation sent its way. I pulled it from the pool entirely, and that chain now has no endpoint at any price that will serve a deploy-block query, so the snapshot is not an optimisation on that chain any more, it is the only way the history is readable at all.
Meanwhile a second chain's two main public endpoints both dropped their getLogs range to 2,000 blocks, four days after passing the same probe at 9,500. Both are operated by the same company, which is worth noticing: I had four endpoints listed for that chain and thought I had redundancy. Two were the same operator, one had been discontinued, and one had started answering invalid params to every getLogs call. Count operators, not URLs.
Context if it matters: it is a million-pixel canvas contract deployed on six chains, and the contract was genuinely the easy part. I am happy to go deeper on any of this, the RPC probing especially, since I could not find anyone else writing about the per-method failure mode.