Skip to content

About PunkRaven

We build the layers underneath.

We build self-hosted AI infrastructure for Indian languages and high-stakes domains: a speech layer across all 22 scheduled Indian languages, and a reasoning layer that grounds every claim in a real retrieved source. Alongside them we are building one application, LawSafe. Both layers are built to run on hardware you control, and to say plainly when they cannot verify what they are about to tell you.

A foreign API comes with someone else's priorities.

India is online at enormous scale, and the great majority of its internet users read, watch and search in an Indic language rather than in English.

The country runs on more languages than almost anywhere on earth. It generates document and audio volume at a scale few countries match. It has a domestic software industry with the talent to work on any of it.

What it does not have is a stack of its own. Almost every Indian product that needs to hear, read or reason reaches for a foreign API. With it come that vendor's language priorities, that vendor's pricing power, and that vendor's willingness to sound certain about things it has no basis for.

The gap is not ambition or talent. It is that the layers underneath were built for somewhere else. The part in the middle - the part that decides whether the thing works in Marathi, or on a body of law that changed last year - is nobody's product, so it becomes everybody's bug.

We build that part. It is slower, it is harder to fund, and it is the part everything else depends on.

A set of build decisions, not a label.

Building for a market is easy to claim and hard to verify. Here is what we mean by it, concretely, in properties you can check.

PunkRaven builds systems for the language rather than translated into it, trained on the real domain rather than approximating it, on open weights you deploy and run yourself with no outbound path. The economics stay predictable, and every system is built to say plainly when it cannot verify an answer. Each is a property you can check, not a label.

Systems that know what they do not know.

There is one belief underneath everything PunkRaven builds, and it is unfashionable: a system that admits uncertainty is more valuable than one that never appears uncertain.

The industry has already demonstrated the alternative, and the clearest published evidence comes from the domain we chose first. Stanford RegLab's Large Legal Fictions study found general-purpose models inventing answers to legal questions often enough that no practitioner could rely on them, and its follow-up found that even purpose-built retrieval-augmented tools still fabricated.

The failure mode is not domain-specific. It is that these systems are wrong fluently, with no signal to the reader that anything has gone missing, and that is as true of a transcript as it is of a citation.

So we design in the opposite direction. A claim either traces to retrieved material or is withheld. Every transcript segment is built to carry a confidence score you can act on. Every language carries an honest quality tier.

Abstention is a designed behaviour, deliberately trained, not a failure mode we apologise for. Verification is built to run in layers and fail closed: if a layer cannot confirm a claim, the system withholds it rather than shipping it unverified.

This costs us things. Our demos are less impressive. Our systems will sometimes say they cannot help. We think that is the correct trade in every domain where a wrong answer is expensive, and those are the only domains we intend to work in.

Two layers, and one application.

PunkRaven is a product company, and the product is a stack. The two infrastructure layers are the company. Alongside them we are building one application of our own, LawSafe. LawMan reasons; LawSafe is what a person opens.

The infrastructure

  • TNT - the language layer

    Status: Planning

    A self-hosted speech pipeline that turns Indian-language audio into a clean transcript and a translation through one API call. Two engines, one queue, one deployment unit.

    The part in the middle - voice activity detection, punctuation, sentence splitting, number formatting, protected-term handling - is the part everyone else leaves to you, and it is where most avoidable quality loss happens. Shipping it as product rather than as a tutorial is the reason TNT exists.

    Its buyers are contact centres, consumer apps, government services, media and education - anyone whose users do not speak English.

  • LawMan - the reasoning layer

    Status: Specified

    A system designed not to answer from memory. It is built to retrieve the governing material first, answer from what it found, attribute the claim to its source, and verify each reference against the actual source text before you see it. Skill lives in the model; facts live in the sources.

    Its first body of authority is Indian law, chosen because it is the least forgiving test available. The sources are authoritative, the language is exact, and a confident invention is not a rough draft but a liability.


Alongside it

  • LawSafe - the first application

    Status: In design

    A chat-first way for any Indian to describe a legal problem in their own language and understand where they stand, in plain terms and grounded in cited sources, then reach a Bar Council-verified advocate who specialises in that issue. Understanding first, transaction second.

    It is a separate application in its own right, built so an ordinary person can understand their legal position, in their own language and grounded in real sources, before spending money on advice.

The constraints define us as much as the features.

  • We will not ship a system that sounds certain when it is not.

    Abstention is the product working, not the product failing.

  • We will not publish a benchmark we have not run.

    Estimates are labelled as estimates, every time, including when it would be more persuasive not to.

  • We will not put a customer logo on this site before there is a customer.

    A trust row of placeholder logos is the fastest way to lose a technical reader.

  • We will not make your data the business model.

    Self-hosting is not an enterprise tier we upsell. It is how the systems are built.

  • We will not represent ourselves as a law firm or a legal services provider.

    PunkRaven builds software. LawMan and LawSafe are research and drafting instruments; advice comes from a qualified advocate, and every surface we build makes that boundary explicit.

  • We will not let commercial pressure override the grounded-or-silent rule.

    If that rule ever becomes negotiable, the rest of this page is marketing.

Pre-launch, and specific about it.

We are a young company doing the unglamorous half of the work first, on the theory that the layers underneath are the only part that is hard to copy.

  • TNTStatus: PlanningA complete technical specification, a costed deployment plan and a documented API contract. At planning stage.
  • LawManStatus: SpecifiedA full technical specification of the retrieval, attribution and reference-verification design. Specified in full, not yet built.
  • LawSafeStatus: In designA product vision and a defined scope, with the understanding flow mapped. In design, with no advocate panel yet.

If you have Indian-language audio, a body of authoritative material that has to be reasoned over without leaving your infrastructure, or simply a reason to keep that work on hardware you control, we would like to talk. That includes the case where the honest answer is that we are not ready for you yet.