90 Day Carrier Performance Scorecard for Ops: 4 TMS Metrics
A pragmatic 90 day plan for carrier operations to build an automated, TMS‑sourced, lane-level performance scorecard using four KPIs to drive corrective...
A carrier performance scorecard is a set of documented KPIs, pulled straight from your TMS and EDI feeds, that tracks how well your own operation executes against your own targets. The fastest route to a working version is four metrics with fixed formulas, scored by lane rather than fleet-wide, refreshed automatically, and reviewed monthly. Skip the spreadsheet and you skip most of the reasons scorecards die within a year. Logivo is one option built to run this natively.
TL;DR:
- Carrier scorecards should focus on four key metrics: OTIF, damage claims ratio, billing accuracy, and carrier utilization or tender acceptance, reviewed monthly.
- Building scorecards from your TMS and EDI data ensures accuracy and reduces disputes, requiring six to twelve months of shipment history for reliable benchmarks.
- The scores should be weighted according to your business priorities, with lane-level scoring revealing performance issues hidden in fleet averages.
- Regular, automated reviews—weekly alerts, monthly operational meetings, and quarterly improvement plans—drive meaningful carrier performance improvements.
- Using an automated platform like Logivo ensures real-time data updates, clear corrective actions, and reduces reliance on manual, error-prone spreadsheet processes.
LogivoBring Clarity to Transport OperationsLogivo brings AI recommendations, delivery tracking, job allocation, and invoicing together in one transport management platform.Explore Logivo
Table of Contents
Four metrics beat fourteen. A short list gets reviewed every month; a long one gets ignored after the second quarter, according to guidance on building carrier scorecards in seven steps. Pick metrics tied directly to outcomes your customers and your margins actually feel.
- On Time In Full (OTIF): the share of shipments delivered both on schedule and complete. You can calculate it by multiplying separate on-time and in-full percentages, or by counting the orders that hit both conditions at once. The second method is more statistically sound, because on-time and in-full failures are rarely independent events, and multiplying inflates the apparent failure rate.
- Damage and claims ratio: claims filed (or units damaged) divided by total shipments over the period, tracked per lane where volume allows.
- Billing and invoice accuracy: invoices matching the agreed rate card without a manual correction, divided by total invoices issued.
- Driver utilisation or tender acceptance: pick whichever reflects your model. Utilisation (active driving hours over paid hours) fits owner-driver fleets; tender acceptance rate suits brokers managing a mixed carrier base.
Whichever OTIF formula you choose, document it once and never switch mid-year, or every trend line you produce becomes meaningless.
Why your TMS and EDI feeds must be the source of truth
Carrier-submitted spreadsheets flatter the carrier. A TMS execution record does not. Building the scorecard from your TMS and EDI data rather than carrier-reported numbers removes the single biggest source of dispute, and it requires six to twelve months of shipment-level history before any baseline means anything statistically.
Map each KPI to one feed before you build anything:
- OTIF — TMS delivery timestamps cross-checked against IFTSTA status messages where EDI is in play.
- Damage and claims — your claims log, reconciled against POD notes and driver defect reports.
- Billing accuracy — invoice line items matched against the TMS rate card at the point of invoicing.
- Tender acceptance or utilisation — tender logs and driver activity records inside the TMS.
Native TMS connectors or a light EDI/ETL pipeline keep these feeds current without a human re-entering numbers weekly. That single decision, according to platform comparisons of carrier scorecard software, is what separates scorecards that survive from ones that quietly stop being updated.
Pro Tip: Set alerts on trend shifts, not every event. A scorecard that pings a manager for each late delivery trains people to ignore it within a fortnight; one that flags a lane sliding from amber to red for two straight weeks gets acted on.
How do you set scoring bands, weights and lane-level scores?
Weighting is where strategy gets encoded into the numbers. A study on scorecard weighting frames this plainly: the weights should reflect what actually matters to your business, not an equal split across metrics because equal felt fair.
A thin-margin bulk hauler might weight billing accuracy and damage ratio more heavily, since disputes eat margin faster than a missed delivery window does.
- Set bands per KPI (green/amber/red), then apply weights to produce one composite score per carrier, per lane using this Lastaufnahmemittel Auswahl checklist for transport managers to optimize operational readiness.
- Score by lane, not only by fleet average. A carrier can post a strong overall OTIF while one lane quietly runs at 70%, and the seven-step build guide is explicit that lane-level scoring is what surfaces this.
- Lock the weights for the full evaluation period to preserve comparability over time.
- Revisit weights once a year, tied to a genuine shift in priorities, never on a whim.
A 7-step plan to build your scorecard in 90 days
Treat this as a build sequence with named owners, not a wish list.
- Define KPIs and formulas. Ops lead documents the four metrics and locks the OTIF calculation method.
- Map each KPI to a data source. TMS administrator confirms which field or EDI transaction feeds each metric.
- Pull historical baseline data. Data or IT lead extracts six to twelve months of shipment-level records per carrier, matching the minimum baseline window recommended for reliable comparisons.
- Build lane-level baselines, not just carrier averages, so early anomalies are visible from day one.
- Automate the data pull and set alert thresholds so scores update without manual entry.
- Pilot on a sample of lanes or carriers before rolling out fleet-wide.
- Formalise review cadence and ownership, naming who owns the data, who owns the conversation with each carrier, and who signs off on any contract action.
Pro Tip: Judge the pilot on three things: does the automated feed match your manual spot checks, do the alerts actually fire when they should, and can the system generate a quarterly review packet without someone building it by hand at midnight?
How often should you review scores and act on them?
A scorecard that never triggers a decision is a report, not a management tool. Scores only change behaviour when they’re tied directly to concrete actions and the cadence is enforced without exception.
- Weekly: automated alerts on lane-level anomalies or sudden score drops.
- Monthly: operational review of the full scorecard against targets, owned by the ops lead.
- Quarterly: a formal business review where sustained amber or red scores trigger a documented improvement plan.
- Annually: an audit that feeds directly into contract renewal, volume allocation, or removal decisions.
Green carriers earn more volume. Amber carriers get a 30 to 60 day improvement plan with named corrective steps. Red carriers move to probation, and repeated red status without improvement is grounds for removal from the panel.
What pitfalls quietly kill a scorecard’s credibility?
Most scorecards fail for the same handful of reasons, and none of them are exotic.
- Tracking too many KPIs. Fourteen metrics nobody checks are worse than four that get reviewed every month.
- Accepting carrier-submitted PDFs as fact, which reopens every dispute the scorecard was meant to close.
- Changing weights mid-period to make a preferred carrier look better, which quietly destroys trust in the whole programme.
- No clear data owner, which is consistently cited as the reason scorecards decline after the first few enthusiastic months.
Run two validation checks quarterly: reconcile your OTIF figures against raw POD and TMS timestamps, and sample a handful of invoices against the rate card to confirm billing accuracy hasn’t drifted. A widening gap between the two is your trigger for an unscheduled audit.
How do industry benchmarks put your scores in context?
There is no single universal OTIF or damage-ratio benchmark that applies across every haulage, trucking, or drayage operation, because acceptable performance shifts with contract type, freight class, and lane distance. A dedicated retail replenishment lane and a long-haul bulk contract simply carry different realistic ceilings for on-time performance.
What holds constant across operation types is the discipline behind the number: a documented formula, a consistent measurement window, and a source of truth that doesn’t change definition halfway through the year. Rather than importing someone else’s benchmark wholesale, build your own internal baseline from six to twelve months of shipment data, then treat that baseline as the standard you’re trying to beat. A carrier or lane consistently below your own trailing average is the real red flag, regardless of what a published industry figure says.
Treat published benchmarks as a sanity check, not gospel. If your OTIF sits well below what similar operations typically report, that’s a reason to investigate your data pipeline before assuming your execution is genuinely worse. Sometimes the gap is a timestamp mismatch, not a service failure.
How should you handle exceptions and outliers in the data?
Not every bad score deserves the same response. A single missed delivery caused by a weather closure is noise; a lane that drifts from green to red over six consecutive weeks is signal, and treating both the same way is how scorecards lose credibility with the very carriers they’re meant to hold accountable.
Build a simple exception rule before you need one: any shipment affected by a documented force-majeure event (severe weather, port congestion, a consignee refusing delivery) gets flagged and excluded from the automated score, but logged separately so the exclusion itself is auditable. Without that separate log, exclusions quietly become a way to flatter underperforming lanes.
For genuine outliers, look at frequency before severity. A carrier posting three consecutive weeks below 80% on the same lane, while every other lane it runs stays green, is telling you something specific about that lane, not that carrier’s overall capability. Cross-reference against your TMS execution logs and IFTSTA data before raising it, because a data feed gap will produce exactly the same visual pattern as a real service failure.
Set a standing rule: any score more than two bands away from a carrier’s trailing three-month average automatically routes to a manual review before it hits the monthly report. That single filter catches most data errors before they become a difficult conversation.
How do scorecard results feed carrier improvement programmes?
A score sitting in a dashboard changes nothing on its own. The research on scorecards that actually change behaviour makes the point directly: the number only matters once it’s wired into a decision someone is accountable for making.
Structure the link between score and action as a formal tier system. Green-tier carriers get priority tender access and, where volume allows, a larger share of new lanes. Amber-tier carriers get a documented improvement plan with two or three specific corrective actions and a 30 to 60 day review point, not a vague conversation about “doing better.” Red-tier carriers move to probation with a hard deadline: hit agreed targets by the next quarterly review or lose the lane.
The improvement plan itself should reference the exact KPI that triggered it. A billing accuracy problem needs a different fix (probably a rate card mismatch or an invoicing process gap) than an OTIF problem caused by chronic late departures. Treating every amber score with the same generic “improve performance” letter is how carriers learn to ignore the programme entirely.
Feed the annual audit results directly into contract renewal conversations. A carrier’s trailing twelve-month composite score should be sitting in front of procurement before any rate negotiation starts, not produced as an afterthought once terms are already agreed.
How do you communicate scorecard results without carriers tuning out?
Numbers without context read as an accusation. Numbers with a comparison, a specific cause, and a next step read as a working conversation. Before any monthly or quarterly review, pull the lane-level detail behind the headline score, because a carrier told “your OTIF dropped” will ask which lane, and “overall” is not an answer that builds trust.
Share the exact formula behind every metric before the first review, not after a dispute. A carrier who doesn’t know whether OTIF counts a two-hour delivery window or a same-day window has a legitimate grievance, and that argument is entirely avoidable if the calculation is documented and shared upfront.
Lead each review with what’s working, briefly, then move straight to the specific lane or metric that needs attention. Vague framing like “overall service has slipped” invites a vague response.
Put the improvement plan in writing with dates attached, and confirm the carrier received it. A verbal warning in a monthly call gets forgotten by both sides within a fortnight; a written plan with a 45-day checkpoint gets tracked by someone.
Making scorecards operational, not decorative
Most carrier scorecards fail for a mundane reason: they run on data someone typed into a spreadsheet once a month, and typed numbers drift, get fudged, and eventually get ignored by everyone including the people who built the thing. A scorecard only earns trust when it runs on execution data the operation actually controls, refreshed automatically rather than assembled under deadline pressure.
The operational payoff shows up in unglamorous places. Fewer invoice disputes because the billing accuracy number matches the rate card automatically. Clearer corrective actions because the lane-level detail is already there when the conversation happens. Faster resolution because nobody’s first move is arguing about whose numbers are right. Platforms with native automation, including Logivo’s guided one-month trial, let a team validate this against their own data before committing to anything.
— Vytautas
How Logivo turns scorecard data into daily action
Logivo is the alternative to spreadsheet-built scorecards and manual monthly rollups: your OTIF, billing accuracy, and lane-level scores update automatically from live TMS execution data, not from a file someone has to remember to update.
Job allocation, delivery tracking, ePOD capture, and invoicing all run inside one platform, which means the four KPIs behind a working scorecard are already being generated as a byproduct of daily operations rather than assembled separately at month-end. Role-based access lets an operations lead see fleet-wide trends while a dispatcher sees only the lanes they’re responsible for, and alerts flag a lane’s slide from green to amber before it becomes a quarterly surprise. Because invoicing runs through the same transportation management platform, billing accuracy scores reconcile against the actual rate card automatically rather than against a static spreadsheet copy.
The guided one-month trial exists precisely for this: run it against a sample of your own lanes and carriers before deciding whether to build your full scorecard programme on it. Start the trial and see whether your first automated lane-level score matches what your gut already tells you.
Sources
FAQ
It’s a set of documented KPIs, sourced from your TMS and EDI feeds rather than carrier-submitted reports, used to track and improve your own operation’s execution against defined targets.
How many KPIs should a carrier scorecard track?
Four is the practical ceiling for a scorecard that gets reviewed every month; longer lists tend to get ignored after the first quarter or two.
Should OTIF be calculated by multiplying on-time and in-full percentages?
Counting orders that satisfy both conditions at once is more statistically robust than multiplying the two percentages separately, since the two events are rarely independent.
How much historical data do you need before scoring carriers?
Aim for six to twelve months of shipment-level data per carrier before treating any baseline as reliable enough to act on.
How often should scores be reviewed?
Weekly automated alerts, a monthly operational review, a quarterly business review, and an annual audit that feeds contract decisions is the cadence that keeps a scorecard tied to real action.
Can a TMS like Logivo automate carrier scorecarding?
Yes. A TMS-native platform such as Logivo pulls OTIF, billing accuracy, and lane-level data directly from execution records, removing the manual spreadsheet work that causes most scorecard programmes to quietly stall.
Recommended