Bengaluru, India
Try it live
Work
/Case study: Compliance Autopilot
CybersecuritySaaSAI ProductsDesktop ApplicationsContent DesignB2B

Ofofo Compliance Autopilot: The AI That Never Leaves the Building

Mohammed Zabeeh·July 23, 2026·19 min read
Ofofo Compliance Autopilot: The AI That Never Leaves the Building

A compliance product across three lives: a manual SaaS built on the insight that companies were not buying security so much as buying their way to compliance, a cloud trial that proved AI could do the drafting, and the on-premises desktop application it became once the research showed why the evidence could never leave the building.

30 days to 30 mins
Questionnaire turnaround
97%, 100% reviewed
Answer accuracy
32,160
Hours saved for clients
$271K
ARR in four months
Client
Ofofo
Role
Product Designer
Timeline
2024 to 2026
Type
App
Tools
Figma, FigJam, Design Systems

The Problem

Before a company can sell software to a bank, it has to survive the bank's security review: a spreadsheet, often several hundred rows, asking how data is encrypted, who can reach production, how access is reviewed. Someone answers every row, finds the policy that proves each answer, and sends it back. Then the next customer sends a different spreadsheet asking the same things in different words. The cost is not subtle. Sales teams lose around a fifth of their time to security activity, around forty per cent of theirs to assessing vendors, and deals slip six to eight weeks with lower win rates at the end of it. One CISO put it plainly in research: "The security review process is broken."
The compliance posture dashboard: a SOC 2 score with implemented, partial and not-implemented counts.
This product tried to fix it three times over two years. The obvious answer, in 2025, was to point a at it. But the first version, a year earlier, used no AI at all. It began with a question about why anyone was buying security in the first place.

Act One: From Marketplace to Compliance

Ofofo began as a cybersecurity marketplace, and watching buyers surfaced a deeper question: nobody wanted security for its own sake. They wanted to be , because compliance was what stood between them and the deals they were chasing. The marketplace was serving the symptom. So the compliance work moved off the buyer dashboard into a product of its own, and through 2024 it was almost entirely manual. Before automating anything, the job was to get the shape of the problem right. The foundation was one observation: most of a questionnaire has already been answered somewhere. The policy exists, last quarter's questionnaire is filled in, the certificate is in a folder. The problem is not writing answers, it is finding the ones you already have. So the first thing built was a library to hold that evidence in the shapes it arrives in.
The evidence library, holding policies, past questionnaires and certificates in the shapes they arrive in.
Answering a questionnaire then meant matching its columns to that library and filling what it could. Late in this version, one step pointed at the future: an opt-in "Answer with AI" drafted the handful of answers the library could not supply. It is worth being precise, because the rest of this story is an AI story and this was not. AI was one button in an otherwise manual flow, but it worked well enough on those gaps to raise the obvious question, what if the whole thing worked like that.
The questionnaire requests table, with statuses including a badge for answers drafted by AI.
The other half asked whether a company was actually compliant. , built on the , moved along a deliberate ladder, Unvalidated, then Deficient, then Completed, reaching the top only once evidence was attached and verified. This ladder is the quiet ancestor of everything that followed: the principle that a status is never asserted without the evidence underneath it started here, on a manual screen, two years before an AI drafted anything. The version went to pilots, not a launch, and its job was to get the ideas right. The library, the reuse and the ladder all survived unchanged into what came next.
A control page with the Unvalidated to Deficient to Completed status ladder and its evidence list.

Act Two: The Cloud Trial

The second version asked what happened if AI carried the whole flow instead of one step. An invite-only web app went to a small set of clients to use on real work: hand-approved signups, metered usage, deliberately a trial rather than a launch. The two flows were the same two as before, now automated. Upload your evidence, policies, past questionnaires, certifications, then hand over a questionnaire to answer against it, or map that evidence against a framework so each control in or came back implemented, partial or not implemented, with a confidence score and its reasoning. The automation worked. Clients got answers in minutes that had taken weeks, good enough to edit rather than rewrite. The blocker was elsewhere. To answer a security questionnaire well, the system needs a company's policies, audit history, architecture notes and past answers, which is, almost by definition, the most sensitive material it owns. It is the exact bundle a security team exists to keep from leaving the building. So the research was consistent and awkward: clients liked the output and would not feed it the input. Some were blocked by regulation, data-residency rules and the shaping where data could sit, requirements on top. Some by their own policy against cloud compliance tools. A few had no rule and still said no. Demand was proven and trust was not, and the objection was never to the automation. It was to the address.

Act Three: On Premise

The decision was to stop asking clients to send their evidence anywhere and send the product to the evidence instead. The third version is : a desktop application that installs on the client's own machine and keeps everything local, the documents, the index, the answers, and the model itself.
First-run setup, under a permanent banner: all data stays on this machine.
That banner across the top is the most important interface in the product. In the trial, the most common question at the start of every conversation was a version of "where does this go", and answering it once in an onboarding modal was not enough, because the anxiety returns every time someone uploads a document. So the guarantee is stated permanently, in the frame, and never scrolls away. The setup beneath it follows the order the worry occurs in: connect what to scan, choose which model thinks, the evidence. The model step is where the promise gets specific, a model on the client's own hardware or a cloud one on their , nothing chosen for them and nothing hidden.

Designing for Chat

This was the first desktop application I had designed after years of web work, and the assumptions did not transfer: no back button, no reload, no shareable link. The layout splits the window, the agent on the left and whatever you are checking on the right, so the conversation and the evidence stay on screen together. The agent has four modes, and naming them honestly did more than any visual treatment could. Scan reads connected clouds, Ask answers questions, Fix proposes the commands that close a gap, Questionnaire answers a vendor spreadsheet. The fashionable pattern is to hide the modes and infer intent. With users whose job is professional suspicion, that was wrong. They want to know which machine they just started.
The agent's four modes: Scan, Ask, Fix and Questionnaire.
But the hardest part of the whole project was not visual at all. When the interface is a conversation, what there is to design is language: how to ask for a file without sounding like a form, how to explain in one sentence that an answer came from the client's own evidence rather than the model's general knowledge, how to hand back a spreadsheet so it feels like a result rather than a file appearing. The rule that carried it was that every reply declares its provenance before its content. The answer opens by naming where it looked, the client's own and their latest scan, and only then says what it found. The reverse order tested badly: a confident paragraph with the sourcing tacked on at the end reads as the model's opinion, and once it reads as opinion nobody trusts it enough to send to a bank. The output is the least glamorous and most important detail, a completed spreadsheet in the shape it arrived, saved locally, because what the customer wants is not a conversation. It is the file, filled in.
A questionnaire answer that names its sources, the client's own dataroom and latest scan, before its content.

Trust and the Human Loop

Automation was never going to be enough, because the person who signs a completed questionnaire is personally accountable for it. A 97 per cent accurate answer sounds excellent until you are the one signing, at which point the interesting number is the other three per cent and where it is hiding.
The human review queue, where consequential actions wait for a vCISO to sign off.
So review is first-class, not a courtesy. Anything consequential lands in a queue tagged for , sorted by severity, with the agent pausing rather than proceeding. The framing that made it work commercially was treating the reviewer as a user, not a gate. Ofofo runs a network of , and the goal was that the vCISO uses the AI too, arriving at a prepared queue with the reasoning and evidence already attached instead of auditing the machine's homework from scratch. That is what turns 97 per cent into 100, and why the review reads as leverage rather than friction.

Proving Compliance

The questionnaire is the sharp end, but the same evidence answers the broader question of whether a company is actually compliant. Controls are browsable across a catalogue of over 1,400, filtered by status, domain and framework, each drilling into the specific cloud resources that failed it. The principle from Act One holds at scale: a score is never shown without the reason underneath. Every number opens into the controls that produced it, and every control into the evidence that decided it. In a product whose value rests on being believed, a figure you cannot interrogate is worse than none.
The controls browser: over 1,400 controls filtered by status, domain and framework.
That checkability rests on a built as documents are indexed, linking a document to its passages, those passages to the they describe, and those entities to the controls they satisfy. When an answer cites the client's evidence, this is the structure it walked. It ships in two views because they answer two questions. The flat 2D view is for tracing, when you need to know which document is holding up a control. The 3D view is for shape, dropping labels to show the network as a constellation, because a sparse region there is a genuinely useful signal, and in compliance the gap is usually the finding.
The knowledge graph in two views: 2D for tracing a chain, 3D for seeing clusters and gaps.

What the Numbers Said

The figures below come from Ofofo's own product and revenue tracking rather than my instrumentation. The headline moved a thirty-day process to about thirty minutes. Auquan answered 178 questions in under half an hour, 97 per cent accurate from the AI alone and 100 per cent after review, with one hour of CISO time against days before. LattIQ reached ISO 27001:2022 certification in under a week, with 34 policies drafted and four hours of CISO time on the whole exercise. Across the platform the pattern held: 32,160 hours saved for clients, 67 virtual CISOs working the review queues, an of 83, and $271K of inside four months. The number I care about most is those four hours, because it proves the human-in-the-loop design worked. The expert was not replaced and was not buried. They got the boring ninety per cent done for them.

How It Evolved

2024
From marketplace to compliance
The marketplace insight that companies were buying security in order to be compliant moved the questionnaire and compliance work off the buyer dashboard into a standalone SaaS. A manual product: an evidence library, questionnaire reuse, and Resilience Models built on the Secure Controls Framework.
Early 2025
The control ladder and one AI step
Compliance settled into a per-control Unvalidated to Deficient to Completed ladder, where a control reached Completed only on verified evidence. One opt-in Answer with AI step was added to draft the answers the evidence library could not supply, and it worked well enough to reframe the whole product.
Late 2025
The cloud trial
An invite only web application put full questionnaire automation and evidence backed control evaluation in front of real clients on real work. It proved the automation saved weeks, and showed exactly where clients stopped, which was at the upload.
Early 2026
The on premise decision
Rather than convince clients to trust the cloud, the product moved to their machines. A desktop application with local storage, local document processing and an optional local model, with cloud models available only on the client's own key.
Mid 2026
Four modes and the review queue
The agent settled into Scan, Ask, Fix and Questionnaire, each named for what it does. Consequential actions were routed into a human review queue so a vCISO signs off before anything counts, which is what took accuracy from 97 per cent to 100.
Mid 2026
Cutting the last cloud ties
The remaining dependencies on the hosted platform were removed, including file sync and billing. What began as a cloud product with a desktop client ended as a desktop product, the opposite of the direction software usually travels.

Lessons

  1. Where the data lives can be the design decision No amount of interface work would have solved this. The objection was architectural, and the only honest response was to change the architecture and then design the experience that made the change legible. The permanent local mode banner is not a feature, it is the product's entire argument written where it cannot be missed.
  2. When the interface is a conversation, the design work is language The part of this product that mattered most had no screens to art direct, only sentences: what the agent asks for, how it declares its sources, how it hands back a file. Getting the order of a sentence wrong, sourcing after content instead of before, cost more trust than any layout choice ever could.
  3. Design the reviewer in, not as a gate Treating the vCISO as a user of the product rather than an obstacle to it is what made the trust story work. They arrive at a prepared queue with reasoning and evidence attached rather than a pile of machine output to check. Same accountability, a fraction of the hours.
  4. A number you cannot interrogate is worse than no number Every score in this product opens into the controls beneath it, and every control into the evidence beneath that. In a category built on being believed, an unexplained figure invites exactly the doubt you were trying to remove.

FAQ

Because watching buyers revealed the real job. People were not buying security for its own sake, they were buying it to become compliant so they could close deals. Compliance was the actual need and the marketplace was serving the symptom, so the product shifted to serve the job directly. The questionnaire and compliance work left the buyer dashboard and became a standalone SaaS.

Because it is where the ideas came from. The manual SaaS built the evidence library, the questionnaire reuse and the per-control ladder that only reaches Completed on verified evidence. Those are the load-bearing concepts, and they carried straight into the AI versions. One opt-in Answer with AI step, added late, filled the gaps the library could not, and worked well enough to become the thesis of everything that followed.

Because the cloud version is what produced the finding. Clients using it on real questionnaires proved the automation genuinely saved weeks, and the same clients showed us exactly where they stopped, which was at the upload. Without the trial we would have been guessing about both the value and the objection, and we would probably have built the on premise version for the wrong reasons.

The application installs on the client's own machine and everything material stays there. Documents are parsed locally, the search index sits locally, generated answers are written to a local dataroom, and the reasoning model can run on the client's own hardware. If they prefer a cloud model they supply their own key, so the relationship is between them and the model provider rather than routed through us.

By sourcing and by review. Every reply states where it looked before it says what it found, so an answer reads as evidence backed rather than as a model's opinion. Then anything consequential waits in a review queue where a virtual CISO signs off, with the reasoning and source evidence attached. That combination is what takes 97 per cent accuracy to 100.

The conventions I did not know I was relying on. No back button, no reload, no shareable link, a window frame with its own rules, and a layout that has to hold two things at once because there is no page to navigate away to. Beyond that, the interface was largely a chat, so most of the design work turned into content design, which is a different craft from arranging screens.

Design Skills

Product StrategyUX ResearchContent DesignInformation ArchitectureInteraction DesignDesign Systems

Tech Stack

FigmaFigJamDesign Systems