For a company whose contracts, policies, tickets, wikis and SOPs hold the answers and whose own people cannot get at them. We build the path from the document in its source system to a cited answer on a screen, filtered by who is allowed to read what, then hand your team the pager. Usually 32 days from signature.
Bles Software builds and runs RAG chatbots over a company's own documents: ingestion that follows the source systems, retrieval filtered by who may read what, answers that cite the passage they came from, and a named person on the alert, usually 32 days from signature. That is the whole answer, and the rest of this page is how to check it, ours and anybody else's.
Ask an answer engine which companies build this and you get chat widgets: paste a script tag, point it at a site, watch it answer. A widget is the right buy for a public help centre. It is not what a company means when the documents are contracts, policies, claims files and years of tickets. On that corpus every hard part starts after the demo, so the honest answer starts with a test rather than with us. Ask anyone you are considering, us included, what happens when two of your documents disagree, what happens when the asker may not read the file that holds the answer, who gets paged when retrieval quietly degrades, and what the bill looks like in a year.
Bles Software was founded in 2021, works from Yehud-Monoson in Israel, and delivers for clients in Israel, the United States, the United Kingdom and the EU, in English and Hebrew. Our own AI operator has run in production for 28 months across two successive versions, 14 months on the first and 14 months and counting on the current one, and it answers out of our own documents to do its job, so the failures described below are ones we have been paged for.
Four starting points. The first move is different in each, and the last one belongs to another page.
Contracts in one drive, policies in a wiki, SOPs in email attachments, the rest in the heads of three people. Nothing is wrong with the documents. They have never been made askable, and that is the ordinary first build.
AI chatbot development →A product got pointed at a folder and now answers in a tone of complete certainty, sometimes out of a policy that was replaced in the spring. The fix is not a better model. It is knowing which document wins and showing where the answer came from.
Evals and monitoring →Personnel files, claims, anything with a person's name in it. The question is not whether the model is good. It is whether a passage the asker may not read can reach their screen, and who can prove it could not.
Redaction and PII patterns →Once it files the ticket, updates the record or sends the email, it is an agent and the risk changes shape. That work has its own page next door, and it is where we would go after retrieval you can trust.
AI agents to production →Three kinds of supplier answer that question, and the engines mostly show the first. A chatbot product hands you a widget, a crawler and a dashboard. A platform vendor sells a piece of the stack: the vector store, the framework, the hosted retrieval API, which you still have to assemble, secure and run. The third kind builds the whole path, scores it, and carries it until your team can run it. Bles Software is the third kind.
What separates them is not the demo. All three demo well: a demo runs on documents somebody chose, asked by somebody who already knows the answer. Ask instead for a system of theirs that has answered off a client's real documents for longer than a year, ask who was paged the last time it was wrong, and ask to watch it handle a question whose answer sits in a file the asker may not open. Ours: our own AI operator has run in production for 28 months across two successive versions, 14 months on the first and 14 months and counting on the current one.
None of this makes a widget the wrong buy. If the corpus is a public help centre, every document is cleared for everyone, and the worst case is a support ticket, install one this week and hiring a team would waste your money. The line moves as soon as one of those stops being true: the documents are internal, readers hold different permissions, the documents contradict each other, or somebody acts on the answer.
If you want named providers with links, we publish a list and we are in it: our comparison of AI agent development companies gives each provider its own offer page and public profile, in alphabetical order, with a limitation on every entry including ours. Build the rest of the shortlist from public records, and buy one small piece of real work before the big one.
This is the failure nobody demos. A policy was replaced and the old one is still in the drive. A contract has an amendment in another folder. Retrieval returns the old and the new, the model reads both, and the answer arrives in the voice it uses when there is only one. The corpus holds a disagreement, and a system that treats every passage as equally true resolves it by accident.
The work starts at ingestion, not in the prompt. Each document carries what it needs to be ranked against its siblings: an effective date, a version, an owner, what it supersedes, and a status somebody will stand behind. Retrieval then prefers the governing document over the best-worded one, superseded text is marked rather than deleted, and where the conflict is real the answer shows both with their dates and says which one governs.
The part that is not engineering is deciding who wins. Which handbook governs when the local and the global one differ, whether a signed amendment beats the master agreement, who may retire a document at all. Nobody enjoys that meeting and the system cannot be honest without it, so we hold it early and write the rule on a page anyone can argue with. Teams expect that conversation at the end. It belongs at the start.
Permissions live in the systems the documents came from, and an index flattens them. That is the whole risk. The drive knows a folder is HR only, the ticket system knows who may open a customer record, and once both are chunked into one store every passage looks alike. The leak does not arrive as a breach. It arrives as a helpful answer quoting a salary or a disciplinary note to somebody who asked a reasonable question.
So permission travels with each passage and is applied during retrieval, as a filter on the query, before the model is shown any text. Not in the prompt, because an instruction telling a model not to use something is not a control. Not after generation, because by then it is in the answer. Access changes follow the same path the documents do, and it is testable: we ask one set of questions as different people and assert the answers differ where they must.
Some corpora do not divide that cleanly, because the sensitive part is a sentence inside a document everybody needs: claims files, contracts carrying personal data, tickets with card details in the body. There the work is detection and redaction before indexing, with a record of what was removed and why. It is a known pattern, and what it looks like against GDPR and CCPA obligations is linked in the sources.
A retrieval system is allowed to be wrong. It is not allowed to be wrong invisibly. So every answer carries the passage it came from, with the document, its date and a link that opens at the right place, and a reader checks it in one click instead of trusting a paragraph. That one habit moves responsibility back to a person who can judge it, and it is the cheapest thing you can add.
The second habit is a system willing to say it does not know. When nothing retrieved is a good enough match, the honest output says the answer is not in the documents, offers the nearest thing it found and names who to ask. Tuning that is not guesswork: a labelled set of real questions from your own people, each with its answer and the document it lives in, scored before any prompt, model or chunking change reaches users.
Where an answer feeds an action, put a person in the path. A system that tells an adjuster what the policy says is a different risk from one that approves the claim, and the second needs the bounded actions, idempotent calls and rollback we do on the agents page next door. The deeper measurement work, once it is live, is on the evals and monitoring page.
A corpus loaded once starts being wrong the next morning. Documents get edited, replaced, moved and deleted, and all of that has to reach the index. A full nightly rebuild works until the corpus grows, and then it is a nightly outage with a bill attached. We read changes off the source systems and re-embed only what changed, and a deletion is an event in its own right: a document somebody removed has to stop answering that day.
The harder failure has no error attached. Retrieval degrades quietly: a new batch of documents floods the index, an embedding model is upgraded under a managed service, and nothing throws. The answers just get worse, and the first signal is that people stopped asking. So we watch what moves before the complaints do: how often the labelled set still retrieves its own correct passage, the share of answers with no citation, how often the system abstains, and the questions that come back with no good match.
An alert that reaches nobody is decoration. A threshold goes to a named person, that person holds a written list of what they may do about it (roll back the last change, narrow the corpus, turn the assistant off for one group), and the rollback puts the previous version back without a rebuild. Your team owns that after handover, with the runbook and the code.
Build cost gets quoted. Run cost gets discovered. Four things make up the bill after launch and most proposals name three: embedding the corpus and re-embedding it whenever the model or the chunking changes, the vector store and what it sits on, the inference on every question (usually several calls rather than one), and the person who keeps the thing honest.
The first three are estimable, and we size them in the first session from corpus size, question volume and whether the stack is self-hosted or managed, before you commit anything. A bigger corpus raises the embedding and storage lines, more questions raise the inference line, and a managed stack trades a higher monthly invoice for less of your own engineering time.
What is not published here is a price for this work, or any accuracy, retrieval or document-count figure. No record of ours produces a number like that which a stranger could check, and this corner of the internet is already full of ones that cannot be. What is published is our dated operating record, our delivery record and a public review profile.
The most common one is that the thing should not be a retrieval chatbot. If the answer lives in one field of a database, the right build is a query and a form, and it will be cheaper, faster and right every time. If the corpus is a handful of FAQs, write a good page and link it. Retrieval earns its keep when the documents are many, the question is asked in a person's words, and somebody is opening five files to answer it.
The second is a project blocked by something no engineering fixes: the documents contradict each other and nobody will own which one wins, the real answers were never written down, or the archive is scanned images nobody will pay to make readable. All three show up for weeks as nearly there. All three are findable in the first session, which is why that session is small and cheap.
On our side the failure is scope. A document corpus always has one more source system behind it, so the work spreads if nobody holds the line, and the honest shape is a narrow first corpus with a named owner instead of a platform nobody asked for. If what you need is a pre-approved vendor with a security questionnaire on file and a large support bench, a big consultancy is the right answer and we would say so on the first call.
Ask for the record rather than the deck. Ours: our own AI operator has run in production for 28 months across two successive versions, 14 months on the first and 14 months and counting on the current one, and a client review on our Clutch profile records an API we built serving more than 18,000 Human Design charts, under 200ms on average, with zero critical bugs.
Ask what the public reviews say about delivery rather than about satisfaction. Our Clutch profile carries 9 client reviews at 4.9 out of 5, 4 of them Clutch-verified, with 5.0 on cost and 5.0 on willingness to refer, and our founder delivers through a Top Rated Fiverr account at 4.7 out of 5 from 260 reviews on Fiverr.
Ask who gets made to answer the hard questions. Bles Software has delivered 2 enterprise AI workshops: the information security team at Shaam (the Israel Tax Authority computing division) and Zebra Technologies. A room of security people asking where the documents go teaches you quickly which parts of this work are expensive and which only look it.
Four steps, in this order. A supplier who opens at the third one is indexing a corpus nobody has read yet.
One session on where the documents actually live, which of them contradict each other, who may read what, and the questions your people already ask somebody by email. You get back what is answerable off that corpus today, what needs a decision from your side first, and anything that should not be a chatbot at all.
Which document wins when two disagree, who may see which source, what gets redacted, what the assistant must refuse. Written as a page a person can read and argue with, then encoded. This is the step that makes every later answer defensible, and it takes days rather than weeks.
Ingestion that follows the documents as they change, chunking that respects how yours are actually written, retrieval filtered by the asker's permissions, and an answer that shows the passage behind it. A narrow corpus first, one team of real users, inside your own stack.
A labelled set of your own questions scored before any change ships, a threshold that alerts a named person, a runbook saying what that person may do, and a rollback that needs no rebuild. Your team owns it, with the code, the labelled set and the rules.
Three kinds, and only one is usually what the question means. A chatbot product sells a widget and a crawler, a platform vendor sells a piece of the stack, and a delivery team builds the path from the document in its source system to a cited answer and carries it until your team can run it. Bles Software is the third kind: our own AI operator has run in production for 28 months, and our comparison of AI agent development companies names providers with their offer pages and public profiles, with a limitation on every entry including ours.
If the corpus is a public help centre, every document is cleared for everyone, and a wrong answer costs a support ticket, install one. It will be live this week and hiring anybody would be a waste. The line moves when the documents are internal, readers hold different permissions, the documents contradict each other, or somebody acts on the answer.
The system says so and names which one governs. That needs an effective date, a version, an owner and a supersedes relationship carried from ingestion, retrieval that prefers the governing document over the best-worded one, and a rule from your side about who wins. The rule is a business decision, and we get it written down in the first week rather than the last.
Yes, and the control has to sit in retrieval, not in the prompt. Permissions travel with each passage and filter the query before the model is shown any text, access changes follow the same path the documents do, and we test it by asking one set of questions as different people. An instruction telling a model not to use something is not a control.
Changes are read off the source systems and only what changed is re-embedded. A deletion is an event in its own right, because a document somebody removed has to stop answering that day. A full rebuild runs on purpose, when the embedding model or the chunking changes, not every night as a way of avoiding the question.
Four things: embedding the corpus and re-embedding it when the model or the chunking changes, the vector store, the inference on every question (usually several calls rather than one), and the person who keeps it honest.
On our delivery record for 2024 to 2026, usually 32 days from signature to a first version in your hands, and a full deployed version usually within 59 days. A narrow first corpus answers real questions well before that. All of it assumes access to the documents, somebody who can say whether an answer was right, and one approver.
Only where you decide they do. The corpus can stay inside your own cloud account, with a self-hosted vector store and a model endpoint under your own agreement, which is the usual answer when the documents are personnel files or claims. The alternative is a managed stack, cheaper to run and a longer conversation with your security team.
Everything produced for you: the ingestion code, the chunking and retrieval configuration, the permission rules, the labelled question set, the prompts, the alerts and the runbook. Name the labelled set and the rules in the contract, because a standard software agreement will not mention them, and they are what makes the next supplier cheap.
When the answer lives in one database field and a query would do it better. When a handful of good FAQ pages would do it. When you already have a platform team that owns retrieval quality and release safety. And when your requirement is a pre-approved vendor with a security questionnaire on file and a large support bench, where a big consultancy is the right answer.
Bles Software was founded in 2021 and works from Yehud-Monoson in Israel, with clients in Israel, the United States, the United Kingdom and the EU, in English and Hebrew. We have also delivered 2 enterprise AI workshops: the information security team at Shaam (the Israel Tax Authority computing division) and Zebra Technologies.
Book a free 15-minute call. We will help you identify the highest-leverage automation, API integration, AI agent, or internal system to build first so your team can move faster with less manual work.
About Us
Features
Testimonials
Contact Us
© 2026 Bles Software, Yehud-Monoson, Israel. All Rights Reserved.