Configuration
There is no global config struct yet (the Python SecurityConfig section is
not ported). Detection is tuned through one flat struct,
guard_core_rs::detect::DetectConfig, passed to every detect call, the
global IP gate through guard_core_engine::ip_gate::IpGateConfig, and the
stateful layer (rate limiting, dynamic bans) through
guard_core_engine::rate_limit::RateLimitConfig and
guard_core_engine::ip_ban::IpBanConfig.
IpGateConfig (whitelist / blacklist / exempt_ips)
The global IP gate is the Rust family's minimal port of the reference
engine's global IP stage. Build it once at startup with
IpGateConfig::new(whitelist, blacklist, exempt_ips), which fails closed on
an invalid entry, and evaluate request IPs with IpGateConfig::evaluate:
| List | Semantics |
|---|---|
whitelist |
Allowlist. When non-empty, every IP it does not match is denied (IP not in whitelist) |
blacklist |
Denylist, consulted when whitelist is empty (IP is blacklisted) |
exempt_ips |
Skip-list for known-friendly automation; sets the skip flag, never a deny path |
Matching semantics are identical for all three lists, mirroring the reference
whitelist matcher: a bare IP or a CIDR range (host bits cleared at parse
time), IPv4-mapped forms matched against their IPv4 canonical form
(::ffff:203.0.113.7 matches 203.0.113.7 and 203.0.113.0/24), and no
cross-family matching.
exempt_ips vs whitelist
exempt_ips is noise reduction for known-friendly automation (monitoring
probes, VPN egress, a partner's server), not immunity: it is noise reduction
for known-friendly automation, not immunity; the blacklist, dynamic bans,
route rules and detection still apply. An exempt match sets the same skip
state a whitelist match sets (IpGateDecision::is_exempt) but never adds a
deny path of its own and never opens the whitelist gate: with a restrictive
whitelist, an exempt IP that is not itself whitelisted is still denied. The
stateful stages below (rate limiting, violation counting, dynamic bans) skip
exactly what the reference skips for a whitelist match
(is_whitelisted || is_exempt) and never skip penetration detection; the
user-agent filter below follows the same rule.
Rate limiting (rate_limit)
The sliding-window rate limiter mirrors the reference engine's rate limiter in its in-memory mode (the reference falls back to this exact store when Redis is off; the Redis-distributed mode is a follow-up).
RateLimitConfig carries the knobs:
| Field | Type | Default | Reference knob |
|---|---|---|---|
enable_rate_limiting |
bool |
true |
enable_rate_limiting |
rate_limit |
u32 |
10 |
rate_limit (requests per window, >= 1) |
rate_limit_window |
u64 |
60 |
rate_limit_window (seconds, >= 1) |
enable_rate_limit_auto_ban |
bool |
false |
enable_rate_limit_auto_ban |
endpoint_rate_limits |
HashMap<String, RateLimitEntry> |
empty | endpoint_rate_limits (per-endpoint (requests, window) tier) |
The config constructor fails closed on a zero limit or window (Python's
pydantic rejects them with ge=1) and on any endpoint_rate_limits entry
with a zero requests or window, the same bound the tier entries
inherit. RateLimitEntry::new and RouteRateLimits::new carry the same
fail-closed validation for the tier overrides.
Counting semantics: one sliding log of request timestamps per
(client IP, scope). Before a request is recorded, every timestamp at or
before now - window is evicted; the pre-recording count decides
(allowed = count < rate_limit), and the block reports
count + 1 (the current request included), exactly the reference's
in-memory formulation. The Redis formulation (allowed = count <= limit
over the post-recording rank) draws the same boundary. A blocked caller
retries after the window (Retry-After: <window seconds>, the reference's
429 Too many requests shape). The window store is an LRU capped at 10 000
keys (_MAX_TRACKED_RATE_LIMIT_KEYS).
Scope: check(ip, None) is the global per-IP window (the default pipeline
tier, every endpoint sharing one budget); check(ip, Some(path)) is the
per-endpoint window keyed by (ip, path) (the reference's
endpoint_path-keyed tier).
Rate-limit tiers (check_tiers)
check_tiers(ip, url_path, route, country_of_ip) runs the reference
pipeline's tier order (RateLimitCheck.check in rate_limit.py, via the
Go port's tiersFor/runTier). Every configured tier records one hit
into its own (ip, path) window (each tier is an independent budget,
exactly as the references count), and the first tier that crosses decides
the outcome with its own window for Retry-After:
- endpoint -
endpoint_rate_limits[url_path](exact path match); - route - the decorator's
rate_limit(window default 60,rate_limit_window or 60); - geo - the decorator's
geo_rate_limits: the resolved country's entry with the"*"fallback (country in limits else "*" in limits). Without a country resolver the tier never applies (the reference'sif not geo_handler: return None), even with a"*"entry; - global - the flat
rate_limit/rate_limit_window, keyed by IP alone.
Without any tier configured the sequence collapses to the global tier:
byte-identical to the pre-tier behavior. RouteRateLimits is the validated
carrier for the decorator tiers (None fields inherit or skip). The
references always key the route and geo tiers by the request path, so a
call without a url_path runs the global tier only.
Dynamic IP bans and the auto-ban engine (ip_ban)
The ban store mirrors the reference IPBanManager in its in-memory mode:
ban_ip(ip, duration, reason)recordsexpiry = now + duration; a duration of zero is rejected (BanError::NonPositiveDuration, the reference raises). Re-banning a live IP overwrites its record.- Durations beyond
LOCAL_CACHE_TTL_CAP_SECONDS(3600) are clamped to it, the references' local-store cap (clampToLocalCap, causenot configured, and the PythonTTLCache(ttl=3600)); only the Redis backend honors longer bans. The defaultauto_ban_durationsits exactly at the cap. is_banned(ip)honors expiry with the Go boundary: an IP is banned whilenow <= expiry; strictly past it the entry reads unbanned and is dropped.- The store is an LRU capped at 10 000 entries (silent overflow, the
references'
localCacheMaxSize/maxsize=10000). - Self-ban refusal: loopback (
127.0.0.0/8,::1/128) and configured trusted-proxy targets returnOk(false)and record nothing - the references' self-DoS guard. Trusted proxies parse fail closed (IpBanManager::with_trusted_proxies). - IPv4-mapped addresses canonicalize to their IPv4 form before any store
key, so
::ffff:203.0.113.7and203.0.113.7share buckets, bans, and counters.
IpBanConfig carries the auto-ban knobs:
| Field | Type | Default | Reference knob |
|---|---|---|---|
enable_ip_banning |
bool |
false |
enable_ip_banning |
auto_ban_threshold |
u32 |
10 |
auto_ban_threshold (>= 1) |
auto_ban_duration |
u64 |
3600 |
auto_ban_duration (seconds, >= 1) |
threat_ban_config |
map of category to ThreatBanEntry { threshold, duration } |
empty | threat_ban_config |
Category keys are validated at config time (fail closed, like
IpGateConfig): a key must be a pattern-table detection category or the
rate_limit pseudo-category (valid_threat_categories()), the reference's
ALL_DETECTION_CATEGORIES | {'rate_limit'} set. Unknown keys are rejected,
as in Python and the TypeScript port.
Violation counting and threshold resolution mirror
_resolve_and_apply_threshold_ban (the TypeScript resolveThresholdBan):
ViolationCounters accumulates per IP per category (LRU capped at 10 000
IPs, _MAX_TRACKED_SUSPICIOUS_IPS; an empty category list records
uncategorized), and IpBanManager::register_violations (or the pure
resolve_threshold_ban) resolves in the reference's order:
- banning disabled: no ban (violations still count, so enabling banning later starts from observed history);
- the first listed category whose
threat_ban_configentry's threshold is met (count >= threshold) bans with that entry's duration and reason"<reason>:<category>"; - otherwise the flat threshold, measured against the total of all counted
categories, bans with
auto_ban_durationand the plain reason; - the self-DoS guard can refuse the resulting ban.
The rate-limit pseudo-category feeds the same engine when
enable_rate_limit_auto_ban is on: a rate-limit crossing counts as one
rate_limit violation, so threat_ban_config["rate_limit"] overrides and
the flat threshold backs it up (reason rate_limit_exceeded), exactly the
reference pipeline's behavior.
DetectConfig
| Field | Type | Corpus default | Notes |
|---|---|---|---|
max_content_length |
usize |
10000 |
detection_max_content_length: the semantic budget; the processed content is truncated to this many code points before analysis, and it bounds truncation |
max_full_scan_bytes |
usize |
262144 |
detection_max_body_inspect_bytes: the preprocessor's full-scan cap; content beyond it is handled by the attack-preserving truncation path |
preserve_attack_patterns |
bool |
true |
detection_preserve_attack_patterns: keep attack-relevant regions intact across the decode pipeline |
semantic_threshold |
f64 |
0.7 |
detection_semantic_threshold: per-attack-type semantic threats fire at or above this probability |
threat_score_threshold |
f64 |
1.0 |
detection_threat_score_threshold: the regex-anomaly weight sum at or above which the verdict is a threat |
Scoring semantics
is_threat is sum(regex weights) >= threat_score_threshold or any semantic
threat. threat_score is min(max(regex anomaly, semantic max), 1.0) when
any threat exists, else 0.0. Semantic threats are emitted per attack type
at or above semantic_threshold, with a suspicious fallback carrying the
overall score when no individual type crosses the line.
Request contexts
detect(content, request_context, config) normalizes the context (the part
before the first :) to one of query_param, header, url_path,
request_body, or unknown. Unknown contexts relax the pattern view
filters; embedded-JSON leaf contexts (a :embedded_json suffix) keep the
suffix for validator scoping.
Per-route detection exclusions (detection_exclusions)
The per-request detection exclusion resolution and the multi-surface
request scan mirror guard_core/_utils/detection_config.py and the
pipeline's detectThreat (via the Go port's detectionexclusions.go). The
decision core is guard_core_engine::detection_exclusions::scan_request;
adapters resolve Some(&DetectionExclusionConfig), Some(&route) per request
(the route standing in for request.state.route_config) and translate the
RequestScanVerdict into the pipeline's ThreatFinding.
Resolution (resolve), the reference _resolve_* semantics:
| Surface | Route None |
Route Some |
|---|---|---|
excluded_detection_headers |
defaults + config | merged: defaults + config + route (additive, never a replacement) |
excluded_detection_params |
config | route replaces |
excluded_detection_body_fields |
config | route replaces |
enabled_detection_categories |
config (None = every category) |
route replaces (an empty set disables every category) |
detection_scan_body |
config (None = true) |
route overrides both directions |
All sets lowercase at resolution time, so matching is case-insensitive.
Scan order (scan_request): URL path (url_path), query params skipping
excluded names (query_param), headers (header; an excluded header is
not skipped outright - it scans with its known-false-positive categories
suppressed, ssrf for the address-carrying proxy headers and for
whole-address-chain values, so a payload in the same header still
detects), then the body surface when scan_body is set (form pairs,
multipart parts, JSON walks, and the blob fallback, each honoring the
excluded body fields). A value that is a threat but whose categories are
all filtered out ends the scan clean - the reference's per-value filter is
terminal, not a reason to keep scanning later values. Regex threats carry
the category; semantic threats carry none (the reference's semantic
payloads have no category key). Mongo-operator JSON keys
($where, $ne, ...) hit from the walk unfiltered by the category set.
Route toggles: detection_enabled(global, route) applies the route's
enable_suspicious_detection over the global enable_penetration_detection
in both directions at request time, and check_applies keeps the
suspicious-activity check alive while the global flag is on or any route
enables its toggle (the reference SuspiciousActivityCheck.applies_to).
Request size and content-type limits (request_limits)
The route-scoped size and content gate mirrors the reference engine's
request_size_content check
(guard_core/core/checks/implementations/request_size_content.py). The
decision core is guard_core_engine::request_limits::decide; the tower stage
is guard_core_rs::request_limits::RequestLimitsStageLayer.
Both limits are route config in the reference (RouteConfig.max_request_size,
RouteConfig.allowed_content_types); there is no global knob, so the Rust
ContentLimits carries both as Option and the all-None default is inert:
| Field | Type | Default | Reference knob |
|---|---|---|---|
max_request_size |
Option<u64> |
None |
RouteConfig.max_request_size (bytes) |
allowed_content_types |
Option<Vec<String>> |
None |
RouteConfig.allowed_content_types |
Behavior, in the reference's order:
- Size first: a missing or empty
content-lengthheader passes; a value at or under the limit passes; one over answers413 "Request too large"(reasonRequest size {n} exceeds limit: {max}). The parse mirrors Python'sint(): surrounding whitespace and a leading sign are accepted, and a negative value passes (-1 <= limit). - Then the media type: the exact prefix before the first
;(no trim, case-sensitive) must be inallowed_content_types, else415 "Unsupported content type". A missingcontent-typeheader compares as the empty string and is blocked by any non-empty allowed list, exactly as in Python. - Fail-secure: a non-integer
content-lengthheader mirrors the reference'sint()ValueError:decidereturnsErr, and the tower service answers the pipeline's fail-secure shape500 "Security check failed"(fail_secure = Truedefault).
The stage learns a path's limits through a RouteLimitsResolver
(path -> Option<ContentLimits>, the tower counterpart of the reference's
request.state.route_config); an unconfigured path passes.
Not mirrored: the reference's EVENT_CONTENT_FILTERED events, log_activity
entries, passive_mode (no Rust config surface yet), the on_block hook,
and custom_error_responses body overrides. The decision core returns the
reference reason strings for adapters that log.
Blocked user agents (user_agent)
The user-agent filter mirrors the reference engine's user_agent check
(guard_core/core/checks/implementations/user_agent.py) and its matcher
(_user_agent_matches_blocked_pattern / is_user_agent_allowed). The
decision core is guard_core_engine::user_agent::UserAgentFilter; the tower
stage is guard_core_rs::user_agent::UserAgentStageLayer.
The config knob is the reference blocked_user_agents list, where every
entry is a regular expression:
| Field | Type | Default | Reference knob |
|---|---|---|---|
UserAgentStageConfig.blocked_user_agents |
UserAgentFilter (compiled blocked_user_agents) |
empty, never blocks | SecurityConfig.blocked_user_agents |
UserAgentStageConfig.ip_ban |
IpBanConfig |
off | enable_ip_banning + thresholds |
Behavior:
- Matching: the
User-Agentheader value is truncated to its first 512 code points (_MAX_USER_AGENT_MATCH_LENGTH) and tested against every pattern with search semantics under the reference compiler's case-insensitive + multiline flags; any match blocks. A missing header reads as the empty string. - Validation:
UserAgentFilter::newfails closed on a pattern the ReDoS validator rejects (_validate_blocked_user_agents_valueraises the same way) and, diverging earlier, on a pattern the regex engine cannot compile (the reference would raise at request time and trip the fail-secure 500). - Skip state: the stage passes for
is_whitelisted \|\| is_exempt, read from anIpGateDecisionrequest extension. - Route lists: the reference tries
route_config.blocked_user_agentsbefore the global list; the builder'sroutesseam resolves a path to a compiled per-route filter (same route seam asrequest_limits). - The block shape:
403 "User-Agent not allowed", answered even without a client IP. - The ban feed: a block runs the reference's
escalate_identity_violationshape - only aThreatFindingrequest extension markedis_threatcounts (empty categories recorduncategorized), feedingregister_violationsunder reasonBlocked user agent: <truncated user agent>with the stage'sIpBanConfig. The 403 goes out either way; the ban answers the next request. The builder'sban_engineseam shares theIpBanManager/ViolationCounterspair with the rate-limit stage (the reference's module-singleton); without it the stage holds a private pair.
Not mirrored: log_activity/event emissions, passive_mode (no Rust config
surface yet), and the log_sensitive_* redaction of the user agent in
reasons.
The tower stage (guard_core_rs::tower)
The facade crate carries the first pipeline stage: a tower::Layer
(RateLimitStageLayer) for Axum/tonic-shaped stacks that wires the stateful
modules above into one request pass, mirroring the reference pipeline's
behavior for the two checks the stage owns (rate_limit, the ban check of
ip_security, and the detection feed of suspicious_activity):
| Request state | Decision |
|---|---|
| no client IP | pass through |
| banned (no exemption skip) | 403 "IP address banned" |
over the rate limit (skipped for is_whitelisted \|\| is_exempt) |
429 "Too many requests" + Retry-After: <window> |
| detection finding crosses a ban threshold (skipped for whitelisted, never exempt) | 403 "IP has been banned" |
| everything else | pass through |
- The client IP comes from the
SocketAddrrequest extension (the peer address), falling back to the leftmostx-forwarded-forentry, thenx-real-ip; a custom extractor replaces the default policy. - The skip state is an
IpGateDecisionrequest extension, exactly what the global IP gate leaves behind; bans and detection still apply to an exempt IP. - A crossing feeds
register_violationswith therate_limitpseudo-category (reasonrate_limit_exceeded) whenenable_rate_limit_auto_banis on; the429still goes out and the ban answers the next request, as in the reference. - A
ThreatFindingrequest extension (what a prior detection stage inserts) feeds the same engine with its categories (reasonpenetration_attempt); the crossing request itself is answered with the 403 crossing-ban shape. - The stage config is the pair of stateful configs above
(
RateLimitStageConfig { rate_limit, ip_ban }); construction fails closed on any invalid part (rate limit bounds, trusted-proxy entries, ban-config validation). - Rate-limit tiers: the service resolves the request path
(
request.uri().path()) and the tier surfaces ride alongside it. The endpoint tier is the config'sendpoint_rate_limitsmap; the decorator tiers arrive either as aRouteRateLimitsrequest extension (what an adapter's routing layer inserts) or through the builder'sroute_resolverseam (path -> Option<RouteRateLimits>, the same route seam the request-limits and user-agent stages use; an explicit extension wins). The geo tier resolves its country through the builder'sgeo_handlerseam (Arc<dyn GeoIpHandler>); without a handler it never applies. A tier crossing answers the same429shape with the blocking tier's window and feeds the auto-ban engine exactly as the global tier does. With no tiers configured the stage is byte-identical to the pre-tier global-only behavior. - Not mirrored: the reference's per-tier event emissions and reason
strings (
EVENT_DECORATOR_VIOLATIONpayloads; the decision names the tier instead),log_activity, and the suspicious-activity400answer for a threat below the ban threshold (the detection stage has no tower counterpart yet).
The actix-web and rocket example stages (examples/)
The same stage ships wired for two more frameworks as example workspace
members, reusing RateLimitStage::decide() as the single decision point so
the family shapes are byte-identical to the tower stage. Both hold no
security logic of their own; the translation from native request types to
decide() inputs and back is the wiring an adapter performs.
examples/actix_app(src/stage.rs): an actix-webTransformmiddleware (RateLimitStageTransform) installed withApp::wraporScope::wrap. The client IP comes from the peer address (req.peer_addr()), falling back to the tower default's forwarded-header policy (leftmostx-forwarded-for, thenx-real-ip) reimplemented over actix's types: actix-web 4 still speakshttp0.2 on its public surface while the facade stage speakshttp1.x, so the tower default extractor cannot be reused verbatim. actix-web's ownConnectionInfo::realip_remote_addrmachinery is not consulted.examples/rocket_app(src/stage.rs): a request guard (RateLimitGuard) plus ignite fairing plus scoped catchers, therocket-guard-rsadapter's pattern. Rocket guards cannot respond directly, so the block answer is stashed in request-local state and the error outcome dispatches to the matching catcher, which renders the family shape. The client IP comes from Rocket's ownclient_ip()seam: the configuredip_header(X-Real-IPby default) when present and parseable, else the remote peer. Mind the posture difference from the tower default (peer first): a direct-exposure deployment should setip_header = ""so a client-supplied header cannot spoof its throttling identity. A request arriving while the managed stage is absent is refused500(fail-secure), never passed uninspected.
In both stages the skip state (is_whitelisted, is_exempt) and the
detection result arrive exactly as in tower (an IpGateDecision and a
ThreatFinding, in request extensions or request-local cache entries
respectively), and the trusted-proxies seam stays on the stage builder
(RateLimitStage::builder(..).trusted_proxies(..)).
What is not implemented (fail-closed honesty)
The port targets spec 4.1.0; the detect stage passes the corpus gate (184 cases, zero xfail). The remaining parity gaps are pipeline-side and listed here rather than hidden:
- Config, pipeline, and handler parity gaps: no Redis-backed
distributed rate-limit mode, no Redis event pipeline, and no
decorator/protocol layer (that is the adapters' job). Present today: the
global IP gate (
whitelist,blacklist,exempt_ipsviaip_gate) and the route decorator'sip_whitelist/ip_blacklistgate (ip_gate::RouteIpGate, thecheck_route_ip_accesssemantics), the in-memory rate limiter and dynamic IP ban store (rate_limit,ip_ban), the rate-limit/ban pipeline stage for tower stacks (guard_core_rs::tower, including thepassive_modeswitch, the endpoint/decorator/geo rate-limit tiers, and the suspicious-activity400contract answer) plus example wirings for actix-web and Rocket (examples/actix_app,examples/rocket_app), the route-scoped request size/content gate (request_limits), the user-agent filter (user_agent), cloud-provider listing (cloud_provider), geo country rules (geo), the security-headers manager (security_headers), CORS (cors), and the response-sideprocess_responsepass with the behavior-rule engine (guard_core_rs::process_response,behavior). The published framework adapters live in the adapter repos. PerformanceMonitorand per-scan timeouts, plus a handful of tracked detection knobs recorded as unmapped with reasons.- Behavior-rule storage is in-memory only: the behavior engine's sliding windows and ban dispatch live in the process-local stores; the reference's Redis-backed layout for behavior rules has no distributed mode in this port.
- Route IP-list order: this port evaluates the route
ip_blacklistfirst, then a configured routeip_whitelisttakes over the route verdict (a miss denies, a match passes); the reference (check_route_ip_access) evaluates the route whitelist first, so a whitelisted IP passes even when also blacklisted. Recorded divergence.
Conformance knobs
The conformance harness maps the five DetectConfig fields above from the
corpus config_knobs and records every unmapped knob with a reason. Changing
engine behavior requires the ledger or xfail baseline to stay consistent:
the gate fails on unbaselined failures, stale xfails, and not-run corpus
cases alike.