Skip to content

Operations

How to run a portal automation trial

The way to evaluate automation for portal work is to run a short trial on your real tasks, then compare three things across the tools you are considering: whether each run actually finished, how exceptions surfaced, and how much operator time it saved once rework is counted. This is the hands-on checklist for doing that.

It assumes a non-technical operator runs the trial with a supervisor reviewing.

Step 1: Pick three to five representative tasks

Use tasks your team actually does every week and that matter financially. Demo tasks will not tell you anything.

Cover different shapes of work, because they fail differently:

  • Read-only. Status checks, pulling reports.
  • Write-heavy. Forms, submissions, uploads.
  • Mixed. Log in, search, update a record, download a file.

Include at least one write task. Web Bench, a benchmark across 2,454 tasks and 452 sites, found that current automations succeed on roughly 46.6 percent of write-heavy tasks against more than 75 percent on read-heavy ones. A trial made only of status checks will flatter every tool in the field.

A realistic five-task set for an insurance agency or a billing team:

1. Log into a carrier or payer portal and pull last month's statement. 2. Check status for a list of claim numbers and download any decision letters. 3. Submit a reconsideration form with five to ten fields filled from a spreadsheet. 4. Upload supporting attachments to a record, then confirm they appear. 5. Log into a supplier or VMS portal, export a CSV, and drop it in a shared folder.

Step 2: Define "completed" in operator terms

For each task, write down four things before you run anything:

  • Inputs. What the automation receives: the claim list, the policy numbers, the date range.
  • Portal steps. The high-level actions: log in, search, open, download, submit.
  • Output. What must exist at the end: a file, an updated record, a confirmation.
  • Proof. How you will verify it: a file in the right folder, a changed status in the portal, a timestamped record.

Worked example, for the reconsideration form. Inputs are the spreadsheet row plus the static values like facility ID. Steps are log in, find the claim, open the form, fill every required field, attach the document, submit. Output is the submission appearing in the portal's own history. Proof is a record carrying the claim ID, the submission date, and the confirmation text the portal returned.

That definition is what separates "the tool ran" from "the work is done." Online-Mind2Web, a live-site benchmark across 300 tasks and 136 sites, had to move to human judgment because automated checks were consistently over-optimistic. Your supervisor's read is the same correction inside your business.

Step 3: Keep a trial log

One row per run, in a spreadsheet: date, portal and task, which tool, operator, start and end time, result (completed, exception, or not started), and a notes field for what had to be fixed by hand.

This is the whole measurement apparatus. It is also the thing that keeps the comparison honest, because it records the runs that went badly rather than the ones anyone remembers.

Step 4: Run every task from login, in every tool

Do not let a tool start mid-session or skip authentication. Logins, sessions and two-step checks are exactly where portal work breaks, so a trial that begins after login has skipped the hard part.

For each run: have the operator describe the task in their own words, use identical inputs across tools, require each tool to handle login through its normal path, and note whether the operator needed technical help to get started. That last observation is data. If setting up a basic portal task needs an engineer, the tool is built for a different buyer than the one running your portals.

Step 5: Measure completion, not run success

After a week or so of runs, sort them into three buckets per tool and per task:

  • Completed. Met your written definition.
  • Exception. Stopped, and said clearly what blocked it.
  • Hidden rework. Reported success but still needed a manual fix.

Counts are enough. You do not need percentages, and you should be suspicious of anyone who offers you theirs.

That third bucket is the one worth the trial. Most automation reports success whether or not the work finished, and the team finds out days later through a missed deadline. Rindler is built to remove that gap: a run either completes or says plainly that it did not, and names what stopped it. Your log is how you check whether that holds on your sites rather than on ours.

Step 6: Watch how exceptions behave

Exception handling is a buying criterion, not a nice-to-have. During the trial you will naturally hit portal layout changes, a new required field or validation rule, an extra step in the login flow, and the occasional unexpected page. When each happens, ask four questions:

  • Did the run stop cleanly, or claim success anyway?
  • Did the operator get a message that explained what blocked it?
  • Could the supervisor find a record of it without hunting?
  • How much work was it to fix and rerun?

What you want is a clear, human-readable explanation, enough detail for an operator to tell whether to correct the data or escalate, and proof of what happened even on the runs that stopped. What you do not want is a log file that only a developer can read. If understanding a failed run requires engineering time, every exception becomes a ticket.

Step 7: Quantify time saved, including rework

For each tool and task, pull three numbers from the log: average minutes to do it by hand, average minutes for automated runs that completed cleanly, and the average minutes added by rework on runs that claimed success but needed fixing. Time saved per run is manual time minus completed-run time, and rework comes off the top.

Counting rework is the part teams skip, and it is where an unreliable tool hides. A run that finishes in seconds but needs a five-minute check every time has not saved five minutes.

Time is not the only stake. HP's 2025 survey found that about 85 percent of workers name repetitive tasks as a top contributor to burnout, and portal work is usually exactly that kind of task. A tool that saves minutes but replaces them with worry has not helped.

Step 8: Decide on four criteria

At the end, map your notes to four things you can explain to anyone in the organization:

1. Operator fit. Can the people who do the work set up and run tasks without technical help? 2. Completion proof. Is it obvious that the work finished on the portal, with records you trust? 3. Exception handling. Does it stop cleanly and tell you what happened? 4. Time saved. Does it cut portal time without pushing debugging onto the team?

Share the summary with your operations leader and whoever owns cost or risk. They should be able to see why the choice is safe, not just why it was impressive in a demo.

FAQ

How long should a trial run?

Long enough to catch normal portal trouble. Two to four weeks across three to five tasks is usually enough for patterns to show.

How many runs per task?

Aim for at least ten per task per tool. That is enough to tell a one-off from a repeat problem. You are not running a statistical test; you are finding out whether operators trust the runs.

What if a tool looks great in demos but struggles on our portal?

Trust your portal. Benchmark results skew optimistic when the tested tasks are narrow, and a demo is the narrowest task there is. Frequent exceptions or hidden rework in your own log is the finding.

Do we need an engineer to evaluate this?

No, and needing one is itself a result. The evaluation belongs to the operator and supervisor who live in the portal every day.

Automate the sites your work depends on.

Tell us the sites and the work. We go through them and get them running.

Start free trial