Approach
Measure the fetch before arguing about the answer.
Why Arcus starts at retrieval rather than at generation, where relevance judgements actually come from, and the list of ways this premise could turn out to be wrong.
First principle
1The order of operations
Fix the input before tuning the thing that reads the input.
A retrieval augmented pipeline has a strict dependency. The generation step can only reason over the passages the retrieval step handed it. That makes the quality of the fetch an upper bound on the quality of everything after it, and it makes tuning the generation step before measuring the fetch a way of spending weeks improving how well a model reasons about the wrong documents.
This is not a controversial claim in information retrieval. It is close to the founding assumption of the field. It is nonetheless routinely ignored in practice, and the reason is not ignorance. It is that the generation step is easy to inspect and the retrieval step is not. Anyone can read an answer and form an opinion. Reading a ranked list of forty passage identifiers and forming an opinion about whether the right one is missing takes a labelled set, and the labelled set does not exist yet.
So the work is not persuading anyone that retrieval matters. Most engineers running these systems already believe it. The work is making the measurement cheap enough that it happens on an ordinary week rather than after an incident.
The hard part
2Where relevance judgements come from
This is the expensive, unglamorous centre of the whole thing, and it is worth being honest about it.
Every measure on the front page needs a set of judgements underneath it. For a given question, which documents in the corpus actually answer it, and how well. Somebody has to decide that, and the somebody has to be a person who understands the domain.
Why it cannot simply be automated away
The obvious shortcut is to ask a language model to judge relevance. It is a real technique, it is used in published work, and it is genuinely useful for expanding a small labelled set. It is not a replacement for human judgement, for a reason that matters here more than usual. If the same family of models is used both to retrieve and to judge, the evaluation inherits the retriever's blind spots and reports back that everything is fine. An instrument built from the thing it is measuring is not an instrument.
Our position is that a small set of human judgements beats a large set of automatic ones, that automatic judgements are a reasonable way to widen a human labelled set once agreement between the two has been measured on a sample, and that the agreement measurement is not optional. That is why label agreement is on the shelf next to the measures rather than hidden in an appendix.
Pooling, and the documents nobody judged
A corpus of two hundred thousand documents cannot be judged exhaustively against three hundred questions. The standard answer, used at the Text Retrieval Conference since the beginning, is pooling. Run several different retrieval configurations, take the union of the top results from each, and judge only that pool. Everything outside the pool is treated as not relevant.
That treatment is a known approximation and it has a known bias. A genuinely relevant document that no configuration in the pool ever retrieved is silently scored as irrelevant, which flatters every configuration equally and hides the same blind spot in all of them. Any honest report has to say how deep the pool was and how many configurations contributed to it, because those two facts bound how much the numbers can be trusted.
If a product ever exists here, the labelling cost stays with you. We can make the judging interface fast, we can sample intelligently, and we can tell you when your judges disagree. We cannot make the reading go away, and any pitch that suggested otherwise would be the tell that the pitch was not written by someone who had done it.
Boundary
3Why the product stops where it stops
The most useful thing a small company can publish is the list of things it has decided not to do.
The pressure on anything in this space is to expand. Measure retrieval, then measure answers, then store the documents so measuring is easier, then serve the retrieval itself since you are already holding the documents, then host a model because the customer asked. Each step is individually reasonable and the destination is a product with no shape.
The boundary here is a single sentence. It reads result lists and it computes numbers. Everything on the far side of that sentence is somebody else's job.
- No generation, ever. Not summaries, not explanations of a score in prose, not a chat interface over the results. A number and a sorted list of the worst questions is the whole output.
- No index and no corpus. Document text would not be requested and would have no field to sit in.
- No hosted retrieval. The pipeline being measured stays yours. An evaluator that also serves the traffic it evaluates is marking its own homework.
- No performance monitoring. Latency and cost are worth measuring and are already well served. Mixing them in here would let a fast pipeline look accurate.
- No claim about model quality. If the passages were right and the answer was wrong, this says so and then goes quiet.
A boundary that is written down is a commitment. A boundary that lives in somebody's judgement is a preference, and preferences move the first time a large enough customer leans on them.
Risk
4Ways this premise could be wrong
Written down while it is still cheap to be wrong about them.
- Context windows keep growing until retrieval stops mattering. If a model can read the whole corpus, ranking is a cost optimisation rather than a correctness problem. We think cost and latency keep retrieval alive at any realistic corpus size, but that is a belief and not a measurement.
- Nobody will pay for the labelling. The measurement is only as good as the judgements, the judgements need domain experts, and domain experts are the most expensive people in the building. A tool that requires a week of their time may simply never be used, however good it is.
- The platforms absorb it. Retrieval evaluation is a natural feature for whoever already sells the index or the model. If it arrives bundled and adequate, a separate tool has to be considerably better than adequate.
- Open source already covers it. There are respectable open libraries that compute these measures. If the gap is genuinely only the harness and the report, that gap may be too small to build a company inside.
- Teams do not actually want to know. A measurement that says the pipeline shipped last quarter retrieves the right document for six questions in ten is unwelcome information. Unwelcome information has a poor adoption record.
None of those are resolved. They are listed because a page that describes only the reasons something will work is marketing, and because the fifth one worries us most.
The company
5Register facts, verifiable at source
Everything below can be checked against a public register rather than taken from us.
- Legal name
- ARCUS AI PTY LTD
- Entity type
- Australian proprietary company
- ACN
- 697 547 505
- ABN
- 82 697 547 505
- GST
- Registered for GST
- State
- New South Wales
- Registered
- 2026. Nothing has been published since
- Trading name
- Arcus is a trading name of ARCUS AI PTY LTD
- Verification
- The ABN, the entity status and the GST position are published free of charge at abr.business.gov.au. The ACN sits on the register maintained by the Australian Securities and Investments Commission at asic.gov.au
- Service of documents
- The registered office recorded against ACN 697 547 505 at ASIC is the address with legal effect for service. We do not publish a second address here, because a second address would not have that effect
The company publishes no postal address on this website. Australian law requires the company name and ACN on public documents under section 153 of the Corporations Act 2001, and it does not require a registered office address to be published on a website. Where an address with legal effect is needed, the ASIC register is the authority and this page points at it rather than duplicating it.
Honesty
6What we will not claim
The standing list, kept on the site rather than in a policy nobody reads.
- No customers are named, because there are none.
- No case study, testimonial, logo wall or user count appears anywhere on this site, for the same reason.
- No benchmark number is quoted, because the prototype has only ever run against a query set we wrote ourselves, and a query set written by the people being measured proves nothing.
- No certification, attestation or audit is claimed. None is held.
- No launch date is given, because we do not know one and inventing one would be the easiest lie on the page.
- No screenshot of a product appears here, because there is no product to screenshot and a mock up presented as software is the oldest trick in this category.
ARCUS AI PTY LTD does not hold ISO/IEC 27001 certification, a SOC 2 Type I or Type II report, an IRAP assessment or an Essential Eight attestation, and will not represent otherwise until one is genuinely held. It has published no benchmark results, no paper and no reproducible experiment. It has no customers, no revenue and no outside investment.
Write to contact@arcusai.fyi if anything on this site reads as a claim we have not earned. That is a defect report and it will be treated as one.