Skip to content

Operations

What to ask for instead of uptime

Every automation contract you will be handed leads with availability. The service was reachable 99-point-something percent of the month, and here is the credit if it was not.

For portal work, that number answers a question you did not ask. Your queue does not care whether the service was reachable. It cares whether the claims got checked. Those are different things, and a month can go perfectly on the first measure while failing completely on the second.

Here is what to put in the agreement instead.

Uptime is not the bottleneck

Availability measures the vendor's own systems. Portal work depends on systems neither of you controls: the payer portal, its sign-in flow, its multi-factor challenge, its habit of changing without notice.

A service can be up for the entire month while a portal quietly stops returning results, and the availability report will be immaculate. The number is true and irrelevant, which is the most expensive kind of metric.

Ask for task completion

The unit that matters is the item, not the interval. So the commitment should read in items.

For a defined queue, what proportion of items is expected to reach a terminal state, meaning completed or explicitly reported as an exception with a reason? Note the wording. An item that stops and says why has reached a terminal state. An item that quietly did not run has not, and it is the second category that costs you money.

Ask for the definition in writing, ask which items are excluded, and ask what happens to the ones that fall outside.

Ask for exception transparency

This is the term that protects you, and it is the one least likely to be offered.

Automation that reports success it did not achieve is worse than automation that stops. If a run says it checked 400 claims and it checked 380, the twenty missing do not appear as an error. They appear weeks later as unpaid claims, and by then the run logs have rolled and nobody can identify which pass dropped them.

So require, in the contract:

- every item that did not complete is named, with the portal and the step it stopped on - exceptions are reported when they happen, not summarised at the end of the month - a run that partially completed is reported as partial rather than as a success with a smaller total - the record is per item, and is retained long enough to reconcile against your own systems

The test is simple. Ask the vendor to show you a run that went wrong. If they can only show you healthy runs, the reporting was built for demos.

Ask about turnaround, at the task level

Availability windows say nothing about how long your queue waits. If the work is checked once a day, a portal that fails at 9am is a day lost regardless of what the availability figure says.

Agree how quickly a queue is picked up, how quickly an exception reaches a human, and who that human is on your side. A fast response to an alert nobody is watching is not a fast response.

Ask what happens when the portal changes

Portals change. That is the recurring event this whole category has to survive, and it is usually absent from the agreement entirely.

Three questions settle it:

- Who notices? If the answer is that you notice when the numbers look wrong, the arrangement has no change management in it. - Who fixes it, and on what timeline, and is that timeline in the contract or in a support queue? - What does the queue do meanwhile? Holding items for a human is a reasonable answer. Marking them complete is not.

Ask who can change a task

Not strictly a service term, but it determines your cost for the life of the agreement.

If changing a workflow means filing a request with the vendor, then every small adjustment carries a lead time and a price, and the long tail of smaller queues never gets automated because each one needs a project. If the person who owns the workflow can describe the change themselves, that cost disappears.

Ask to see someone who is not an engineer make a change during the evaluation.

A short checklist

Take this into the next vendor conversation.

1. What proportion of items reaches a terminal state, and how is that defined? 2. Show me a run that failed, as your reporting displays it. 3. Are exceptions reported per item, when they happen? 4. How long until a queue is picked up, and until an exception reaches a person? 5. When a portal changes, who notices, who fixes it, and what happens to the queue meanwhile? 6. Can a non-engineer change a task while I watch?

None of these are exotic asks. They are what you already assume you are buying, and availability language quietly substitutes for all six.

The point of the exercise is not to extract a harsher commitment from a vendor. It is that the work gets done, and that you can prove it went through.

Automate the sites your work depends on.

Tell us the sites and the work. We go through them and get them running.

Start free trial