How a Private On-Premise AI Deployment Works: Black Lily

Black Lily installs an open-weight language model on hardware your firm owns, together with the software your staff uses it through: a chat window, document upload, and search across everything you have loaded. It works the way the AI assistants your team already knows work. The difference is that nothing they type and no file they open leaves your network. Pull the network cable and it still answers.

Why it runs on your hardware

What is on the box

The model is the part everyone asks about. What your team actually works in is everything around it.

Connectors included: Slack, Box, Dropbox, Microsoft 365, Google Workspace, Salesforce, SEC EDGAR. Your documents and your prompts stay inside the network; the connectors and the optional market data feed pull in rather than push out.

Who we build these for

How we put one in

  1. Scope. A 30-minute call, then a walkthrough of where your documents actually live, which systems should be connected, and which recurring tasks are worth handing to a model. You leave knowing what the system would do on day one, what it would not, and what the machine to run it costs.
  2. Install. We size and configure the hardware, deploy the model and the interface onto it, wire up the connectors you named, index your existing document stores in place, and set up users and permissions. It sits on your network, behind the access controls and the backup regime you already run.
  3. Train. Hands-on sessions with the people who will use it daily, plus written documentation and recorded walkthroughs so a new analyst can be brought up in an afternoon rather than scheduled into a course.
  4. Maintain. Monthly monitoring, model updates as stronger open-weight releases ship, connector upkeep, and adjustments as your workflows change. You are not left owning a box that nobody in the building knows how to service.

Typical delivery is four to eight weeks from scoping to daily use, with hardware lead time usually the longest single item.

The questions we get asked

Which models do you actually deploy?
Open-weight models, chosen at scoping against the hardware you are willing to buy and the work you want done. We name the specific model and its size in the scoping document before anything is ordered, and we tell you what it is weaker at as well as what it is good at. We do not deploy a model whose weights we cannot put on your machine, because that would defeat the entire point.
What does the hardware cost?
A single workstation-class machine with a current-generation GPU covers a small office. Cost scales with how many people query it at once and how much text you want it to hold in context, not with how many documents you load. We size it against your actual concurrency and give you the number and the part list before you commit to anything.
Where do the documents actually live?
Where they live now. We index the stores you name and connect to the systems you already run, which means the box reads from your file server, your Microsoft 365 or Google Workspace tenant, Slack, Box, Dropbox, or Salesforce, and keeps its search index on the same machine as the model. We do not require you to migrate anything into a new repository, and no copy of your documents leaves the network.
What happens when the internet goes down?
The model runs on your hardware, so the AI keeps answering and everything already indexed stays searchable. What pauses is what depends on someone else's network anyway: the connectors stop pulling new Slack messages, new files, and new EDGAR filings, and live market data stops updating. All of it catches up when the connection is back.
Can we get to it from outside the office?
The default is on-site only, because it is the simplest thing to explain to a compliance officer and the hardest thing to get wrong. Remote access over infrastructure you already control is a decision we make together during scoping, alongside who gets it and under what conditions, rather than something switched on by default.
Where do the live stock prices come from?
A market data vendor, over a one-way inbound connection, priced as a separate add-on. Quotes come in. Your documents, your prompts, and your positions do not go out. If you would rather the box had no market data feed at all, that is a supported configuration and we will say so in the scoping document.
Can we keep a record of what staff asked it?
Yes, and for most firms that is a requirement rather than a feature. Prompts and responses are logged on the deployment, under the same named user accounts you configure at install, in a form your own retention and supervision tooling can pick up. What you retain and for how long is your policy to set, and we configure to it.
What happens when a better model ships?
We swap it. Open-weight releases improve on a timescale of months, and the maintenance engagement covers evaluating each one against your workflows and upgrading when it is genuinely better rather than merely newer. The hardware usually outlives several model generations.
Can it connect to our portfolio or order management system?
For reading documents and reference data out of a system, usually yes, and we scope it explicitly. For writing into a trading or accounting system, no. This is a research, document, and drafting tool, and we are not going to put a language model in a path where it can move a position or a payment. If someone offers you that, ask them who signs off when it is wrong.

About Black Lily

Black Lily installs private, on-premise AI systems from Philadelphia, Pennsylvania: a complete language model running on hardware the client owns, for hedge funds, registered investment advisers, and family offices whose most valuable documents are the ones they are not allowed to send to an outside service. Founded by William Dorman, Founder & CEO.

Phone: (432) 234-3779 · Book a free 30-minute scoping call: cal.com/black-lily/30min

Black Lily home · Hedge funds · RIAs · Family offices · Blog · Privacy Policy · Terms of Service