ShellvoideShellvoide

Offensive security, shown in full.

The Shellvoide security blog: AI-driven penetration testing, vulnerability research, and application security insights from the team behind KLUE.

Latest·

Two Code-Execution CVEs in Hydra: CVE-2026-106441 and CVE-2026-106442

Hydra is the configuration framework a large part of the Python machine-learning world uses to wire up experiments and apps. We pointed KLUE, our AI security agent, at it, and it found two ways to run code on a machine that loads an untrusted Hydra config. Both get around the blocklist Hydra had already added to stop exactly this. One slips through gaps in that blocklist; the other goes through the logging system, a door the blocklist never guarded. Both were reported, accepted, assigned CVEs (CVE-2026-106441 and CVE-2026-106442), rated High, and fixed.

kluecase-studyvulnerability-research
Read more

More writeups

16 posts

How KLUE Found an Axios Bug That Can Turn Your GET Into a DELETE

axios is the HTTP client most JavaScript and Node apps use to make requests. While reading through it, our security agent KLUE found a bug in how axios chooses the request method. On the default axios object, if another flaw in your app has polluted Object.prototype, a call you wrote as a safe GET can quietly go out as a DELETE, POST, PUT, or PATCH. Here is how KLUE found it, why only some ways of calling axios are affected, and how it was fixed.

kluecase-study

Three High-Severity CVEs in kin-openapi, the OpenAPI Library Behind Many Go APIs

We asked KLUE, our AI security agent, to review kin-openapi, a library that a large share of Go web APIs use to read and check incoming requests. It found three separate bugs that each let anyone crash the service with a single request. The pattern behind all three is the same: the library already had the right safety check, it was just missing on one of the paths that needed it. All three were reported, accepted, assigned CVEs (CVE-2026-100352, CVE-2026-100353, CVE-2026-100354), rated High, and fixed by the maintainers.

kluecase-study

KLUE vs. a Hardened Target: One Bug, No Signature Required

A detailed bug bounty writeup. Our autonomous agent KLUE spent hours proving a hardened crypto-mining marketplace was hardened, then reverse-engineered its HMAC signing scheme, mapped a STOMP WebSocket buried in the authenticated app bundles, and discovered it performed no authentication at all, routing private per-tenant channels off a client-supplied parameter. An unauthenticated, cross-tenant data leak, responsibly disclosed and rewarded.

kluecase-study

Low, Medium, Critical: How Four Frontier Models Graded the Same Live Target

We ran KLUE against the same live production target four times, once each behind GLM 5.2, DeepSeek V4 Pro, Kimi K2.7 and Opus 4.8, all on an identical thirty minute budget. Same target, same clock, four risk verdicts from Low to Critical. A field report on what each model sees, what it walks past, whether it got the severity right, and why the cheapest run came back with the most.

ai-pentestingklue

Six CVEs in 163 Minutes: An Autonomous Pentest, Reasoning Included

We gave our autonomous pentesting agent, KLUE, one real engagement against a mature open-source codebase and 2 hours 43 minutes on the clock. It came back with six CVE-assigned vulnerabilities, including a chained second-order SQL injection and a template injection that reached a shell. This is the annotated reasoning trace.

kluecase-study

Crawling Isn't Attacking: A Live-Fire DAST Benchmark

We pointed OWASP ZAP, Burp Suite, Acunetix, and our own KLUE at the same live, deliberately-broken web app. Three of them crawled it and reported headers. One forged an admin token, dumped every user, and reset the admin password, all from two requests, unauthenticated. A field report on what dynamic scanners can and cannot reach.

dastdynamic-analysis

Four SAST Tools, One Broken App, and the AI That Hit Zero False Positives

A SAST benchmark on OWASP Juice Shop pitting Semgrep, SonarQube, and Snyk Code against KLUE, our autonomous source code analysis platform. The story is in the precision column, and in the vulnerability classes that have no pattern to match at all.

saststatic-analysis

Could Kimi K2.6 Hold Its Own Against Claude Opus on Real Pentest Work?

We ran four frontier models through the same autonomous pentest engagement. Recall, time to finish, and dollars per run all tell different stories, and Kimi K2.6 turned out to be the surprise on the leaderboard.

ai-pentestingklue

Want a run like this against your own stack?

Powered by KLUE and a certified team, Shellvoide finds the security gaps across your apps, cloud, and systems, and delivers full penetration tests in hours, not days, so you can fix what matters before anyone else finds it.