Repo only had Dependabot security updates (reactive). Add weekly grouped
version updates for Go modules and GitHub Actions to keep dependencies
current and reduce CVE accumulation.
CodeQL go/cookie-secure-not-set (CWE-614) flagged the nz-csrf cookie as
missing the Secure attribute. Mirror writeOauth2StateCookie and derive
Secure from the request scheme instead of hardcoding it: forcing
Secure=true would make browsers drop the cookie on plain-HTTP intranet
deployments, breaking the double-submit CSRF pair and 403-ing every
unsafe request.
Completes the EnableShowInService->HideForGuest rename; the public stats
filter was missed in the previous commit, leaving service/singleton
referencing the removed field and breaking the build.
Document in the tool description and protocol comment that the agent
kills the entire process group/JobObject on return or timeout, so
background jobs must fully detach (setsid/screen/tmux/systemd-run).
Mirror the server HideForGuest flag on services: rename the field
(dropping the gorm default), invert userCanViewService so a service is
visible to guests by default and hidden only when HideForGuest is set,
and flip the public-stats filter in servicesentinel accordingly. Update
the visibility and permission-matrix tests to the new semantics.
compareSemver fell back to lexical string ordering when the numeric
segments were equal, so an agent reporting "2.1.0" compared as older than
the gate constant "v2.1.0" (0x32 < 0x76). This wrongly flagged every agent
as unsupported_agent, breaking server.exec and fs.* on all online servers.
Drop the fallback: semverParts already strips the v prefix, so equal
numeric segments mean equal versions.
The error was further masked because tool error responses shipped a
structuredContent {error_code,error} that violates the tool outputSchema
(which requires exit_code/stdout/...). Strict MCP clients validated it and
rejected the whole response with -32602, hiding the real isError text.
Omit structuredContent on error; the cause stays in content[].text.
Carve a new nezha:inventory:{read,delete,*} resource family out of
nezha:server:*. Listing and deleting servers/server-groups (GET /server,
/server-group, /ws/server, batch-delete/server[-group], MCP server.list)
now require nezha:inventory:*, while nezha:server:* covers per-server
runtime operations (get, exec, fs, config, metrics, batch-move,
force-update).
Also declare MCP tool OutputSchema for exec/fs/meta/server/transfer and
wrap server.list output in {servers,count} so strict MCP clients accept
the structured result. Correct the /file all-of scope entry and the stale
server:read documentation.
The Role field used json:"role,omitempty". Admin is RoleAdmin = 0, so an
admin profile serialized without a `role` key. The admin-frontend gates the
admin menu (user management, settings) on `role === 0`, and its normalizeRole
helper defaults a missing role to non-admin, so admins lost the admin menu.
Drop omitempty so role 0 is always sent. Add a regression test.
Upstream hamster1963/nezha-dash-v2 v2.0.4 reviewed commit-by-commit for
supply-chain risk (dependency bumps only, no new packages/postinstall
scripts/network calls/obfuscation) and mirrored to nezhahq/user-frontend-backup.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
gosec G101 flags apiTokenCtxKey, apiTokenLastUsedCtxKey, and the
ScopeNotificationGroupWrite identifier as hardcoded credentials. They are
context-key names and a scope string, not secrets. Annotate them with
#nosec G101 (matching the existing JWTSecretEnvKey precedent) so the
gosec CI step passes without disabling the rule globally.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
In ConfigUsePeerIP mode ctxWithRealIP validated connectingIp but never
assigned it to ip, so CtxKeyRealIP was left empty. model.CheckIP and
model.BlockIP short-circuit on an empty IP, which silently disabled the
gRPC WAF block table and brute-force token counters for every peer-IP
deployment — including the stream interceptors this path now backs.
Assign ip = connectingIp so the WAF and BlockIP observe the source. The
empty-header path is unchanged (it intentionally opts out of IP WAF).
Add rpc_test.go covering peer-IP IPv4/IPv6, the no-peer error, and the
no-header opt-out so the behaviours cannot be re-coupled.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
- ServeRPC now applies real-IP + WAF interceptors to streaming RPCs
(RequestTask, IOStream), not just unary calls; without them
authHandler.check saw an empty real IP so brute-force BlockIP counters
never keyed on a source and the WAF block table was bypassed at the
stream entrypoint. Factored real-IP resolution into ctxWithRealIP shared
by unary and stream paths.
- handleFsRead rejects negative offset/length; handleFsWrite validates
if_match_sha256 is 64 hex chars before dispatch.
- createAPIToken always verifies each server_id exists (even for admins),
so an admin cannot bind a PAT to a not-yet-created server id that a
future server would auto-inherit.
- Add relay zero-length-chunk rejection regression test.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
- consumeTransferToken now recomputes and constant-time compares the
entry's HMAC-SHA256; previously the signature was minted but never
checked, so the documented "防篡改" guarantee was hollow and security
rested solely on the sync.Map key's randomness.
- CSRF switches from a raw-random double-submit cookie to a signed token
(nonce.HMAC-SHA256 keyed by JWTSecretKey). The middleware now also
validates the signature, defeating sibling-subdomain cookie tossing
where a naive header==cookie pair would otherwise pass.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
GHSA-x6fg-52vr-hj4w: NAT is a commonHandler any authenticated member can
create, and newHTTPandGRPCMux matches r.Host against NAT before dispatching
the dashboard/gRPC handlers, so a member could register the dashboard's own
host as a NAT Domain and hijack global routing. Add IsReservedDashboardHost
(InstallHost/ListenHost plus an operator-declared ReservedHosts list for
reverse-proxy deployments), reject reserved hosts on create/update, and drop
pre-planted records when building the startup cache.
GHSA-39g2-8x68-pmx8: bind-time CheckPermission let a member pre-bind a DDNS
profile ID that the victim would only create later. GetDDNSProvidersFromProfiles
now re-validates ownership by server owner UID and skips foreign-owned
profiles (UserID==0 is treated as a migration artifact, not an admin grant).
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
Introduce Personal Access Tokens (nzp_*) as a stateless auth path alongside
JWT, gated per-endpoint by a scope middleware (nezha:{resource}:{verb}) with
fail-closed empty-scope defaults and a server-id whitelist. Self-management
endpoints (profile, api-tokens, oauth2 bind, refresh-token) explicitly reject
PATs to block privilege-escalation chains. A revoke registry tears down active
long-lived connections (terminal, fm, ws, transfer, mcp) the moment a PAT is
deleted, with a tombstone closing the revoke->register race.
Add an MCP endpoint that proxies tool calls (exec, fs read/write/delete,
transfer) to agents over gRPC, guarded by origin/DNS-rebinding checks, a
per-token rate limiter, audit logging, and a kill switch. Serialize all
sends through the IOStream wrapper to honour grpc-go's concurrency contract.
Add CSRF double-submit protection on unsafe cookie-authenticated methods,
exempting authenticated PAT requests by context identity (not a forgeable
Authorization header). Apply visibility/whitelist filtering consistently
across list, get-by-id, and mutate paths to enforce tenant isolation.
Migrate legacy mcp:* scopes: rewrite read/exec to nezha:* equivalents and
drop dangerous write/delete/wildcard grants.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
Gosec G101 treats the environment variable name NZ_JWTSECRETKEY as a hardcoded secret value. Keep the env-first JWT secret behavior unchanged and annotate the constant so CI only suppresses this false positive.
Co-authored-by: naiba/CloudCode <hi+cloudcode@nai.ba>
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
listServerGroup returned every group, including ones whose only
servers are HideForGuest=true (or that have zero members), to
unauthenticated callers. The group name itself is then leaked to
guests even though every server it references is hidden from them.
Guest visitors now only see groups that contain at least one guest-
visible server. Authenticated members and admins keep full visibility
(including their own empty groups) so management UI still works.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
GET /api/v1/file?id=<server> created an FM stream on the agent stream
and committed real state change (TaskTypeFM dispatched). With JWT
cookie SameSite=Lax a victim's browser would still send the cookie on
a top-level cross-site GET, so an attacker could trick a logged-in
user into opening an FM session on any of their own servers, consuming
resources and triggering the agent's FM machinery without consent.
Mirror the GHSA-8qhj-4f8c-j8qg fix: move the route to POST. SameSite=
Lax cookies are not sent on cross-site POST. Frontend (admin-frontend)
adjusted in a follow-up commit.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
GHSA-vrmh-5mmx-hjwx: GET /api/v1/server/:id/service and
GET /api/v1/service/:id/history both iterate the raw service list and
emit ServiceName / timing for any service that happens to monitor the
queried server, ignoring the owner's EnableShowInService=false flag.
Both routes are on optionalAuth, so unauthenticated visitors could
enumerate hidden services by name and timing.
Introduce userCanViewService(c, service):
- EnableShowInService=true -> always visible
- admin -> always visible
- authenticated owner -> visible (HasPermission)
- everyone else -> hidden
getServiceHistory rejects unknown-or-invisible service with the same
'service not found' message so the endpoint cannot be used as an
oracle. listServerServices pre-filters the sorted service list.
Owners and admins keep their existing visibility into their own hidden
services.
Tests:
- TestUserCanViewServiceVisibleServiceIsPublic
- TestUserCanViewServiceHiddenServiceRejectsGuest
- TestUserCanViewServiceHiddenServiceRejectsForeignMember
- TestUserCanViewServiceHiddenServiceAllowsOwner
- TestUserCanViewServiceHiddenServiceAllowsAdmin
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
GHSA-8qhj-4f8c-j8qg: GET /api/v1/cron/:id/manual changed shared state
on the agent stream, and the JWT cookie is SameSite=Lax so a victim's
browser would send the cookie on a top-level cross-site GET. An
attacker could trick a logged-in user into firing any of their own
cron commands.
Switch the route to POST: SameSite=Lax cookies are not sent on cross-
site POST, closing the CSRF window without introducing a new token.
Tests:
- TestCronManualTriggerRejectsCrossSiteGET locks in that GET no longer
resolves.
- TestCronManualTriggerAcceptsSameSitePOST locks in the legitimate POST
still works.
Frontend (admin-frontend) updated in a follow-up commit.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
Replace the {user_id, ip} claim pair with {keyId, uid}:
- keyId is a 32-byte random id that points to a row in the new
jwt_sessions table holding the real user id, bound IP, UA hash,
TokenVersion and expiry.
- uid is the user id encoded through pkg/idcodec; mismatch between
claim uid and session.UserID trips WAF block on the caller IP.
- identityHandler now rejects unknown/revoked/expired sessions, IP
drift and stale TokenVersion. Refresh updates session.ExpiresAt.
User.TokenVersion bumps on password change and revokes outstanding
sessions, so a leaked JWT secret alone is no longer enough to forge
a token. JWTSession rows are GC'd every 10 minutes (expired + grace
or revoked >24h). OAuth2 callback shares the same issue path.
Includes regression tests for happy path, mismatched claim uid,
revoked session, TokenVersion bump, IP drift and unknown keyId.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
New pkg/idcodec wraps sqids with an alphabet derived from the JWT
secret via HKDF-SHA256 (info="nezha/idcodec/alphabet/v1"). Rotating
NZ_JWTSECRETKEY automatically reshuffles the alphabet, which doubles
as a kill switch for outstanding hashids without touching the encoder
itself.
The base alphabet drops visually-confusable characters (0/O/o/I/l/1)
and MinLength=8 hides small integer ids. Decode round-trips through
Encode to reject inputs that decode by accident under the same
alphabet.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
NZ_JWTSECRETKEY takes priority over config.yaml, env-injected keys are
not persisted and skip the version-driven rotation. When the secret is
auto-generated as a last-resort fallback, write it via a single-field
YAML patch so we never round-trip the in-memory key back to disk.
The JWTSecretKey field is marked json:"-" yaml:"-" so Save() can no
longer leak it. Update ReadEnvFile test to lock in env-over-yaml
precedence.
Co-authored-by: cloudcode <cloudcode@users.noreply.github.com>
NewServerTransferClass merged Verified and acked-rollback rows and sorted
them by AckedAt only. On Windows time.Now() granularity is ~15.6ms, so
MarkVerified and an immediately-following MarkRevertDelivered routinely
share a timestamp. The stable sort then left the Verified candidate
ahead of the rollback that is actually on disk and the agent was locked
out on the next restart. Add transferID as a deterministic tiebreaker so
the later rotation always wins.
- Set GinJWTMiddleware.SigningAlgorithm to "HS256" explicitly so a future
library default change (or an alg:none confusion attempt) cannot weaken
token validation. This matches the current gin-jwt default, so behaviour
is unchanged.
- Set CookieSameSite to Lax: same as the modern-browser default, but
pinned so server-side intent is clear and CSRF on cross-site POST is
blocked while top-level GET (OAuth callback) still works.
- Move the nz-o2s OAuth2 state cookie into writeOauth2StateCookie and
set HttpOnly=true. The frontend does not read this cookie, so HttpOnly
is strictly an XSS-hardening win with no behaviour change.
JWT Cookie HttpOnly/Secure are intentionally left default for now: the
frontend reads \`!!document.cookie\` to display login state and many
deployments terminate TLS at an upstream proxy — flipping those would
require a coordinated frontend change.
Co-authored-by: naiba/CloudCode <hi+cloudcode@nai.ba>
Auth.Check resolved client_uuid → server.ID via UUIDToID without
checking that the resolved server's UserID matched the user that the
secret was bound to. An agent presenting one user's secret could
target a different user's server UUID and impersonate it — poisoning
monitoring state, triggering alerts, or quietly receiving tasks the
real owner expected.
Add authorizeAgentForUUID: if the UUID is unknown we still allow new
registration bound to the secret owner; if the UUID is known but
points to someone else's server we reject. The rejection message
"client UUID does not belong to the agent secret owner" also helps
operators trace which user's secret has leaked.
Note: this changes runtime behaviour after batch-move/server — the
new owner must reconfigure agents with their own secret. That is the
correct contract; the previous behaviour was a cross-user reporting
bug.
Co-authored-by: naiba/CloudCode <hi+cloudcode@nai.ba>
The inline magic check expressed the *invalid* form as
\`byte0 != 0xff && byte1 != 0x05 && byte2 != 0xff && byte3 == 0x05\`,
relying on && to detect a four-byte mismatch. Because && short-circuits,
any payload whose byte0 happened to be 0xff was treated as a valid magic
even if the remaining bytes did not match — almost every random payload
slipped through and only the stream-UUID layer above stood between a
caller with a valid agent secret and a live IOStream session.
Extract the check into isValidIOStreamMagic stated positively (all four
bytes must match) so short-circuit reasoning cannot reintroduce the bug.
Co-authored-by: naiba/CloudCode <hi+cloudcode@nai.ba>
The previous loop only checked ownership inside the
\`server != nil && server.TaskStream != nil\` branch. A foreign online
server returned permission denied for the whole batch while foreign
offline / unknown IDs silently went into the Offline bucket — the
response-shape delta let a RoleMember enumerate other users' online
server IDs by submitting them in batches.
Drop both foreign and unknown IDs silently into the Offline bucket
(without dispatching the upgrade task) so the response is indistinguishable
across those three states.
Co-authored-by: naiba/CloudCode <hi+cloudcode@nai.ba>
createTerminal and createFM correctly check server ownership before
issuing a stream UUID, but terminalStream and fmStream only verified
that the UUID existed. Any authenticated user holding a valid stream
UUID could attach to it, gaining the original creator's live shell or
file-manager session — and the UUID is exposed via URL path (referer
leaks, access logs, browser history, frontend error reporters).
Bind the creator user ID into ioStreamContext at CreateStream time,
expose StreamOwnership and IsStreamAuthorizedForUser, and check
ownership in terminalStream/fmStream before the WebSocket upgrade so a
rejected attempt does not tear down the legitimate stream via defer.
NAT streams are also routed through CreateStream(_, 0); they are not
reachable from /ws/terminal or /ws/file so a sentinel user ID is fine.
Co-authored-by: naiba/CloudCode <hi+cloudcode@nai.ba>
GHSA-6x26-5727-rrm9: a low-privilege member could point a DDNS webhook
at internal or loopback hosts and the dashboard would dial them with the
unrestricted utils.HttpClient.
Extract the notification SSRF defenses (CIDR blocklist, IP-pin DialContext,
SNI preservation, redirect rejection) into reusable helpers in pkg/utils
(NewRestrictedHTTPClient / ResolveAllowedHTTPURL / buildRestrictedHTTPClient)
and route the DDNS webhook through the same path. Replace the notification
inline implementation with a thin wrapper to keep behaviour identical.
Side improvements collected by the refactor:
- prepareRequest now resolves DNS once and returns the paired client, so
the dialer's pinned IP and the validated URL stay in sync (no more
double resolution between prepareRequest and SetRecords).
- response body is drained and closed.
- HttpClient / HttpClientSkipTlsVerify are explicitly tagged unsafe for
attacker-controlled URLs.
Tests cover: hermetic SNI preservation, redirect rejection, dial pin to
the vetted IP, the full blocked-CIDR list at the webhook entry point,
and the verifyTLS↔skipVerifyTLS inversion in the notification wrapper.
Co-authored-by: naiba/CloudCode <hi+cloudcode@nai.ba>