STANDALONE
Standalone Queryless for PortalJS
This document outlines the next iteration of QuerylessAI for data portals: a standalone implementation intended to replace the OpenClaw-based prototype before production deployment for a client.
Pricing is based on a PoC built with this standalone-style approach rather than the original OpenClaw setup.
Why move away from OpenClaw
The OpenClaw-based prototype was useful for getting the first version working quickly, but it introduced major problems for a production-facing Queryless implementation.
Problems with OpenClaw
High costs
OpenClaw is designed for general-purpose agents that can perform many kinds of actions on the host machine, such as reading files, listing directories, and running commands.
That broad capability comes with a large hardcoded system prompt, costing roughly 20k tokens before Queryless-specific context is even considered.
As a result, even simple requests became much more expensive than they should be, and token usage grew very rapidly across longer conversations.
In one benchmark, we tested a 7-message conversation covering common QuerylessAI-PortalJS tasks such as:
- searching for a dataset
- asking a question about the data
- requesting a visualization
The first message was already around 25k tokens. By the seventh message, the conversation had grown to roughly 1M tokens. Each new message made the conversation more expensive because the full conversation history had to be sent back to the model again.
This cost profile would make Queryless difficult to price reliably and unnecessarily expensive even at a basic usage level.
Lower prompt adherence with smaller context-window models
As conversation context grows, LLMs generally become more prone to drift and less reliable at following instructions.
In the OpenClaw setup, the hardcoded prompt added a large amount of context that was mostly irrelevant to the Queryless use case. That increased cost and also made it harder to get reliable behavior from cheaper models such as Gemma 4.
In practice, with Gemma 4, the model would regularly fail to behave as expected. For example, when a user asked to find a dataset, the model would sometimes respond as if it did not know what it was supposed to do.
Insecure by nature
Any LLM-based system is vulnerable to prompt injection or jailbreak attempts.
With OpenClaw, the risk surface was larger because the agent was running inside a real VM with broad execution capabilities. In the worst case, a malicious prompt could try to get the model to:
- execute a binary or command that grants attackers access to the VM
- modify its own context files
- persist malicious behavior across sessions
- generate or distribute malicious links to other users
That level of access was acceptable for prototyping, but it is not a good foundation for a production deployment.
Proposed alternative
We built a separate PoC of Queryless for data portals without OpenClaw.
Instead of relying on OpenClaw as the execution layer, this version calls LLM APIs directly and builds its own context, tool-calling logic, and system behavior around the Queryless use case.
The result was similar output quality with much better operational characteristics:
- Responses were similar in quality
- Gemma 4 became much more reliable, reducing the need for more expensive models
- Costs dropped by 10x or more
- Around
$0.001USD for a simple request, versus around$0.01USD with OpenClaw - It took roughly 30 messages involving question answering and visualization generation to reach 256k tokens, which is within the Gemma 4 context window
- Around
- There was no model access to files or shell execution in the way OpenClaw allowed
- Jailbreaks had less persistence and a smaller blast radius
- We gained lower-level control over the system prompt and product-specific customizations
Based on these cost findings, we defined Queryless pricing with a large safety buffer while still keeping it affordable for clients.
What this standalone version should become
The next step is to turn this lower-cost approach into a production-ready implementation.
Core characteristics of the target architecture:
- Direct LLM API integration rather than OpenClaw
- Explicit control over prompts, tools, and context assembly
- Product-scoped capabilities instead of general VM-level agent access
- Safer execution model with a much smaller attack surface
- Cost profile that supports realistic pricing and production usage
Suggested framework
LangChain is a reasonable candidate for this implementation because it supports:
- agent-style orchestration
- tool calling
- model-agnostic integration
It would provide a similar development shape to the current prototype while giving us more control over cost, reliability, and security.