Why fake ROI spreads so easily
Email and coordination are expensive in the aggregate. Microsoft's 2025 Work Trend Index describes frequent interruptions during core hours and a gap between the capacity leaders seek and the time and energy many workers report. A related Microsoft analysis of the infinite workday shows work activity spreading outside conventional hours.
Asana's 2024 State of Work Innovation reports a large burden from organisational friction and hidden work. Its work-about-work research describes time spent coordinating, searching and managing around the task rather than doing the skilled part.
These are real signals from recent workplace research. They are also easy to misuse. A marketer can take a broad average, multiply it by headcount, assume that software removes a fixed percentage and announce a return. The arithmetic looks precise because the spreadsheet has decimal places. The underlying claim may have almost no connection to the buyer's desk.
A freight mailbox has its own composition. Some messages create work. Some provide awareness. Some are noise. Actionable messages vary from a simple missing-field request to a complex exception. People have different loaded costs. Time released does not automatically become cost removed. Risk and service improvements may be valuable but hard to price.
Finance has seen optimistic transformations before. Operations has seen tools that move effort rather than remove it. The answer is not to stop sizing. It is to use ranges, label assumptions and improve the estimate as evidence moves closer to the desk.
Labour is one lens, not the whole case
Released hours and full-time-equivalent proxies are useful because finance can discuss them. They connect observed volume, handling time and loaded labour cost. They can support a range without pretending every released minute becomes cash.
Labour is not the whole value of a governed work queue. Better visibility can help supervisors see exceptions earlier. Explicit ownership can reduce duplicate handling. Source-linked work and approval history can improve reconstruction. Structured intake can make later system integration easier. Consistent process can reduce dependence on one person's memory.
Those outcomes should be named, measured where possible and kept separate from labour arithmetic. Do not quietly convert every qualitative benefit into money. A risk event that did not happen is difficult to price without a defensible frequency and loss model. Customer trust does not become revenue merely because a slide says so.
Lead with the operating case: the desk needs a more legible and controlled way to turn inbound mail into work. Use labour math as one measurable slice. Track service and control outcomes in their own units, such as queue ageing, duplicate actions, corrections, escalations or reconstruction time.
This keeps a positive case for InboxOS without overselling it. The product is designed to structure inbound freight work, assist preparation and require human approval before every customer reply. Capacity may be released if those steps reduce repeated manual effort. The amount belongs to the buyer's evidence, not a universal vendor percentage.
Start with three buckets
A simple model begins by separating inbound messages into three buckets.
Actionable work requires the desk to decide, update, request, prepare or respond. Examples include bookings, document questions, schedule changes, holds, quotes and invoice disputes. This bucket usually has the highest handling time, but it also has the widest variation.
Awareness deserves a skim or may add context, but does not create a next action for this desk. Examples include copied milestones, partner updates and replies where another party has completed the task. Awareness is not worthless. The estimate should not pretend it can all disappear.
Noise does not need operator attention in the normal flow. It can include promotions, irrelevant circulars, duplicates or messages clearly outside the desk. Even here, a cautious model allows for review and classification errors rather than claiming zero effort.
For each bucket, estimate volume, current handling time, addressable share and realisation. Addressable share asks what portion of the current effort the proposed process can reasonably affect. Realisation asks how much of that theoretical release will appear in practice after review, exceptions, corrections and adoption.
Keep those two ideas separate. A system may address most sorting in a bucket but realise less capacity because operators still verify uncertain items. That verification may be the correct control, especially for customer-facing freight work.
A worked example with assumed numbers
This example is illustrative. Every number below is assumed, not measured customer evidence and not an InboxOS performance claim.
Suppose a freight desk receives 1,000 inbound messages on a working day. For a first-pass model, the buyer assumes that 35 percent are actionable work, 45 percent are awareness and 20 percent are noise. That gives 350 actionable messages, 450 awareness messages and 200 noise messages.
The buyer then assumes current average handling of six minutes for actionable mail, one minute for awareness and half a minute for noise. The daily baseline is therefore 2,100 minutes for actionable work, 450 minutes for awareness and 100 minutes for noise. Total assumed effort is 2,650 minutes, or about 44.2 hours per working day.
Now apply an addressable share, still assumed. Perhaps the proposed workflow can affect 40 percent of actionable handling, 50 percent of awareness handling and 70 percent of noise handling. Theoretical addressable effort would be 840 minutes, 225 minutes and 70 minutes respectively, or 1,135 minutes in total.
Do not call those 18.9 hours "saved." The process still has human review, uncertain items, corrections and adoption friction. Apply a realisation factor. If the illustrative expected case uses 60 percent realisation, estimated released capacity is 681 minutes, or about 11.4 hours per day. A conservative case at 40 percent realisation gives about 7.6 hours. An upside case at 75 percent gives about 14.2 hours.
The range is not a forecast until its inputs are grounded. It is a transparent conversation. A buyer can challenge any assumption: perhaps only 20 percent of mail is actionable, perhaps handle time is lower, or perhaps awareness skims are already nearly free. The model changes visibly instead of hiding disagreement inside one ROI percentage.
Annualising requires more assumptions. The company must choose working days, loaded labour cost and the share of released capacity that becomes economically usable. If the desk uses the time to absorb growth or improve service, that is capacity value, not necessarily payroll reduction. If headcount cost actually changes, finance should document the mechanism and timing.
The worked example also leaves out implementation cost, software cost, integration effort, training, change management and ongoing review. A real business case includes them. The purpose of the three buckets is not to make benefits look large. It is to make the reasoning inspectable.
Use evidence grades that improve over time
An estimate should carry an evidence grade so the reader knows how close the inputs are to reality.
Grade C: assumed. Inputs come from buyer estimates, workshop guesses or clearly labelled illustrative defaults. This grade is suitable for an early conversation. It should produce a broad range and a list of facts to collect, not an investment promise.
Grade B: sampled. Inputs come from a defined sample of the buyer's mail and process. The team observes composition and may time handling for selected work. Sampling method, period, exclusions and uncertainty remain visible. A holiday week, peak season or one unusual customer can distort the result, so the grade is stronger but not final.
Grade A: observed in operation. Inputs come from a controlled pilot or production measurement on the buyer's desk. The organisation can compare current and changed handling, account for review and corrections, and see whether released time is usable. Even Grade A is local evidence, not a universal claim for another company.
The labels are not a formal accounting standard. They are a communication discipline. A Grade C estimate should never quietly appear later as a measured saving. A sampled result should identify how the sample was chosen. A pilot result should state what changed at the same time, because staffing, seasonality and volume mix can affect the comparison.
Move from C to B through discovery. Move from B to A through a bounded pilot with agreed measures. If evidence contradicts the original case, update the case. The ability to revise is a sign that the model is doing its job.
What research justifies, and what it does not
Recent research justifies the belief that coordination, communication and fragmented work consume substantial attention. It supports asking whether better intake and ownership can release capacity. It does not tell a buyer the automation rate for a freight desk.
Microsoft's 2025 Work Trend Index provides workplace context about interruptions and capacity pressure. It does not provide a benchmark for minutes saved by InboxOS. Asana's 2024 research supports the broad importance of organisational friction. It does not prove that every minute of work about work can be removed.
McKinsey's 2024 digital logistics survey reports broad digital investment and positive perceived value among respondents, while also discussing fragmentation, integration, data quality and change management. Treat this as evidence that digital logistics can be worthwhile and difficult. Do not copy a survey-level result into a desk-level return.
The research helps frame the problem and the need for careful implementation. The buyer's own sample and pilot must supply the commercial evidence. This separation is important because external research often measures different populations, definitions and outcomes from the proposed project.
What not to claim
Do not claim that AI saves a universal percentage of freight email effort. Mail composition, work types, current process, operator skill and control requirements differ too much.
Do not turn released capacity into headcount reduction without a documented operating decision. A desk can use released time to absorb volume, reduce backlog, improve service, train people or avoid future hiring. Those outcomes may matter, but they are not the same financial event.
Do not assign a monetary value to every avoided risk. If the company has incident frequency and loss data, it may build a risk model with appropriate review. Without that evidence, say the process improves control and measure leading indicators rather than inventing expected loss.
Do not present broad workplace research as product validation. Microsoft and Asana describe general work patterns. McKinsey describes survey findings in digital logistics. None of those sources tested InboxOS or guarantees its results.
Do not imply that human approval disappears from the cost. In the current InboxOS posture, every customer reply requires human approval. Review time belongs in the model. If the interface makes review more efficient, measure it in the pilot.
Do not hide implementation and change costs. Discovery, configuration, integration, training, security review and operational adoption consume effort. Include known costs and range uncertain ones.
Do not combine labour, service, control and risk into one score that cannot be audited. Keep the units visible. Hours are hours. Queue ageing is time. Corrections are counts or rates. Customer outcomes need their own measures.
Finally, do not call an illustrative example a customer case. Assumed numbers can be useful when labelled. They become misleading when the label disappears.
Estimate in conversation, measure on the desk
A website or early sales conversation can offer a Grade C range to decide whether discovery is worth doing. The output should show assumptions and invite correction. It should not display a mysterious ROI score.
Discovery improves composition and handling assumptions from a real sample. A pilot then tests whether the workflow changes effort and control in practice. Runtime evidence can track work types, corrections, review events and queue behaviour, subject to appropriate data governance and metric definitions.
Finance and operations should agree in advance how released capacity will be interpreted. If volume grows while staffing stays flat, the company may have absorbed growth. If backlog falls, service may have improved. If neither occurs, theoretical release may not have become operational value. Measurement should be willing to say so.
InboxOS follows this commercial posture: use labour-proxy estimates to structure the conversation, then rely on discovery and pilot evidence for a decision. The product remains freight first, single-tenant in the current model and human-approved for every customer reply. None of those facts supplies a universal return. They define the operating design whose value the buyer can test.
Honest sizing is not timid. It gives a sponsor a case that can survive contact with the desk. Assumptions are visible. Ranges absorb uncertainty. Evidence grades improve as the project advances. Benefits remain separated by type. When the measured result is good, the organisation can trust it. When it is weaker than expected, the organisation can adjust before scaling.
Sources
- Microsoft Work Trend Index 2025 (PDF): interruption load and capacity gap.
- Microsoft, Breaking down the infinite workday (2025).
- Asana State of Work Innovation 2024.
- McKinsey, Digital logistics: Into the express lane? (Dec 2024).
See the product posture
InboxOS turns operational email into structured work with human approval before every customer reply. Start on the customer page, then request a private sample demo.