AI API Security: Blocking Automated Abuse at the IP Layer
Every call to your LLM endpoint costs real money — in compute, tokens, and API credits. Unlike a traditional web page, a malicious actor hammering your AI endpoint with junk requests can cost hundreds of dollars before your rate limiter catches up. IP-layer filtering is the cheapest control you can deploy.
Why AI endpoints are uniquely attractive targets
Traditional rate limiting assumes traffic originates from browsers. AI API consumers are headless — no cookies, no JavaScript fingerprint, no browser quirks. A simple cURL loop and a stolen API key is all an attacker needs. The payoff can be substantial: free tokens, data exfiltration via prompt manipulation, or burning a competitor's usage quota.
The limits of rate limiting alone
Rate limits are per-key or per-IP but rarely both simultaneously. Botnets side-step per-IP limits by distributing requests across thousands of exit nodes. If your AI gateway only rate-limits by API key, a single compromised key routed through a botnet can exhaust your quota in seconds. IP reputation adds a second, independent dimension of control.
Adding IPAbuse as a pre-flight middleware
Place the reputation check before your model invocation. Reject requests from IPs with a score below your threshold before spending a single token. The latency cost is under 50 ms on cached IPs — negligible compared to a model round-trip.
import { NextRequest, NextResponse } from "next/server";
const IPABUSE_API_KEY = process.env.IPABUSE_API_KEY!;
const RISK_THRESHOLD = 60; // block IPs scoring below this
export async function ipReputationMiddleware(req: NextRequest) {
const ip =
req.headers.get("cf-connecting-ip") ??
req.headers.get("x-real-ip") ??
req.headers.get("x-forwarded-for")?.split(",")[0].trim() ??
"unknown";
if (ip === "unknown" || ip.startsWith("127.") || ip.startsWith("::1")) {
return null; // allow loopback
}
const res = await fetch(
`https://api.ipabuse.org/v1/ip/${ip}/reputation`,
{ headers: { "X-API-Key": IPABUSE_API_KEY }, next: { revalidate: 300 } }
);
if (!res.ok) return null; // fail open — don't block on API errors
const { data } = await res.json();
if (data.reputationScore < RISK_THRESHOLD) {
return NextResponse.json(
{ error: "Access denied." },
{ status: 403 }
);
}
return null; // allow
}Wiring the middleware into your AI route
Call the check at the top of your route handler, before any token is consumed. This pattern works with any model provider — OpenAI, Anthropic, Google, or a self-hosted inference server.
import { streamText } from "ai";
import { openai } from "@ai-sdk/openai";
import { ipReputationMiddleware } from "@/middleware/ipReputation";
import type { NextRequest } from "next/server";
export async function POST(req: NextRequest) {
// 1. Block bad IPs before spending any tokens
const block = await ipReputationMiddleware(req);
if (block) return block;
// 2. Normal AI route logic continues here
const { messages } = await req.json();
const result = streamText({
model: openai("gpt-4o"),
messages,
});
return result.toDataStreamResponse();
}Cache results to stay within rate limits
IPAbuse caches reputation scores on its end, but adding an in-process cache eliminates the network hop entirely for repeat visitors. A 5-minute TTL is a good default — abusive IPs rarely become clean that quickly.
const cache = new Map<string, { score: number; ts: number }>();
const TTL_MS = 5 * 60 * 1000; // 5 minutes
export async function getCachedReputation(ip: string): Promise<number> {
const hit = cache.get(ip);
if (hit && Date.now() - hit.ts < TTL_MS) return hit.score;
const res = await fetch(
`https://api.ipabuse.org/v1/ip/${ip}/reputation`,
{ headers: { "X-API-Key": process.env.IPABUSE_API_KEY! } }
);
const { data } = await res.json();
cache.set(ip, { score: data.reputationScore, ts: Date.now() });
return data.reputationScore;
}Setting the right threshold
A threshold of 50 blocks the majority of known bad IPs with minimal false positives. For high-value or authenticated AI endpoints, consider tightening to 70. For public demo endpoints, a looser 30–40 range avoids blocking curious developers who share IPs with a misbehaving neighbour. Monitor your 403 rate to tune accordingly.