These are research artifacts and open-source tools published by FlowX.AI, not shipped platform features. For platform capabilities, see AI in FlowX.
border: open-source LLM guardrails
border is an open-source Python library (Apache-2.0) that checks LLM inputs and outputs and signs an evidence record for each check. Evidence records hold hashes, the resolved policy hash, and the model revision that produced each finding, never the user’s text. It runs on CPU with the network interface down, and every classifier is scored in 26 languages.border.flowx.ai
Project site and documentation
GitHub
Source, Apache-2.0
Benchmarks
Published latency and accuracy numbers
Open models on Hugging Face
The flowxai organization on Hugging Face publishes 50 Apache-2.0 models focused on on-device AI for regulated industries: PII detection (piiguard, cee-pii), scam and fraud classification (scam-guard), content moderation, and regulated-advice detection, alongside two public benchmarks (cee-pii-bench, scamguardbench).
huggingface.co/flowxai
All models and benchmark datasets
Paper series
Seven technical papers on enterprise AI agents, published at flowx.ai/research:flowx.ai/research
The paper series, with arXiv links where available
Browser-agent reliability benchmark
Browser Agents Don’t Fail on Capability. They Fail on Consistency benchmarks nine LLMs driving the open-source browser-use library through nine controlled web tasks, five repetitions each: 405 runs in total. Tasks cover multi-page forms, iframes, shadow DOM elements, scrolling, and error recovery. Every outcome is verified against the target site’s recorded submissions and state changes, not against the agent’s own success claim. Key findings:- Several models passed all 45 of their runs, and the models that did span a 26.5x per-run cost difference: price does not predict reliability.
- Some models reported success even on runs that failed, with overconfidence rates reaching 26.7%.
- Checking only whether the requested action happened misses over-actions, such as an agent submitting the same claim twice; outcome-state verification catches them.
Read the benchmark
Full methodology and per-model results on the FlowX.AI blog

