Two overlapping scopes answer different questions
Hermes-attributed usage: 6,186 calls in 384 retained usage rows, August 6–September 29, 2026. Three compression calls recovered from Monitoring bring priced coverage to 6,189 calls. Two further logged compression calls have no usable token receipt. Three other Gemini calls have no provider or endpoint and remain excluded from confirmed Vertex usage.
Entire accessible cloud project: 21,751,087,441 recorded tokens across 20 model identifiers (19 with nonzero tokens), October 30, 2025–September 29, 2026. This includes traffic from other applications; it is not proven to be Pavel's personal or Hermes-only consumption. It includes the local calls. Never add these two scopes.
All seven projects visible to the existing cloud credential were queried from creation to the cutoff. Only one returned usage. No returned data does not establish zero lifetime spend, and inaccessible projects remain outside coverage.
The usable Hermes cost estimate is conditional—not money proved paid
Using the published global Standard PayGo rates, correct cache buckets, and advertised promotional credits, the two output-accounting scenarios are US$108.71 and US$116.39. Exact values: $108.714270075 / $116.388642825. Both include $0.12264675 for the three temporal recoveries.[1]
- Scenario A: price stored completion tokens only. This can omit thinking tokens.
- Scenario B: price stored completion plus separately stored reasoning. This can double count if a receipt already included thoughts.
- Before advertised credits: the same scenarios are US$189.20 / US$201.88. Do not add pre-credit and promotional figures together.
- Actual billed and paid amount: unknown. These scenarios are neither an invoice nor lower/upper bounds. They omit unknown calls, unrecorded older usage, non-token charges, taxes, and any account-specific adjustments.
- Standard tier is an explicit assumption for local receipts, which have no retained service-tier field. The broader project has Priority traffic. The audit uses published prices as of September 29, not verified historical invoice SKUs.
All four locally attributed models are included
| Model | Calls | Scenario A | Scenario B |
|---|---|---|---|
| gemini-3.5-flash | 624 | $28.23 | $30.90 |
| gemini-3.6-flash | 17 | $0.25 | $0.27 |
| gemini-3.7-flash | 2 | $0.00 | $0.01 |
| gemini-3.8-flash | 5,543 | $80.10 | $85.09 |
| Local subtotal | 6,186 | $108.59 | $116.27 |
| Additional temporal recoveries: 3.8 Flash | 3 | $0.12 | $0.12 |
| Combined priced coverage | 6,189 | $108.71 | $116.39 |
Displayed lines round individually; exact CSV/JSON values are used for totals. The two unpriced calls are not included in these amounts.
Main chat—not compression—dominates local use
| Task | Calls with local usage | Scenario A | Scenario B |
|---|---|---|---|
| main agent loop | 5,877 | $104.76 | $112.04 |
| background review | 137 | $3.13 | $3.31 |
| compression | 21 | $0.63 | $0.63 |
| approval | 106 | $0.05 | $0.19 |
| title generation | 44 | $0.02 | $0.09 |
| vision | 1 | $0.01 | $0.01 |
The compression row is the local 21-call subset. Its three temporal recoveries add $0.12264675, giving the earlier 24-call $0.75180375 estimate. Two logged calls remain unpriced. The $0.75 is already contained in the combined Hermes estimate; do not add it again.
Cache and reasoning accounting change the answer
Retained local counters: 77,707,943 uncached input, 424,526,358 cache-read, 0 cache-write, 1,139,081 stored output, and 1,631,146 separately recorded reasoning tokens. Source-code inspection proves local input is already net of cache. The reconstructed prompt volume is 502,234,301 tokens; subtracting cache again would undercount.
The historical Vertex OpenAI-compatible route stores completion tokens unchanged and reasoning separately. In 291 of 384 rows, reasoning exceeds output, including 67 single-call rows. Therefore a blanket claim that reasoning is already included is contradicted by these receipts. The native Gemini adapter uses a different route and cannot prove the historical Vertex accounting contract. The two scenarios expose that ambiguity instead of hiding it.
For each model, the calculation is (uncached input × input rate + cache reads × cache rate + stored output × output rate) / 1,000,000; Scenario B adds reasoning × output rate / 1,000,000. Cache-write is zero in these records.
Global Standard rates per million: Gemini 3.5 Flash $1.50 input / $0.15 cached / $9.00 output; 3.6, 3.7 and 3.8 Flash $0.75 / $0.075 / $3.75 during the promotion. Output pricing covers response and reasoning. Google describes the promotion as 50% credits back; actual application is unverified.[1]
The local database's historical estimated-cost sum, $92.830173375, is retained only as an unverified old estimate. It is not the newly calculated result or an invoice. Zero actual-cost fields without billing provenance were treated as unknown.
The project-wide inventory is much larger—and not personally attributable
The independently checked Monitoring inventory contains 17,565,500,860 input and 4,185,586,581 output tokens, plus 1,982,469 invocations across all response codes, of which 1,971,056 returned HTTP 200. Failed/status-diverse invocations are not all billable calls.
Monitoring's token counter is an accumulated input/output count, not a cost ledger.[2] The returned series lack cache quantities, modality, and per-request context tiers. They also include dedicated capacity, unspecified shared tiers, and legacy Priority labels. Pricing all input as uncached Standard or dedicated usage as PayGo would fabricate a project dollar total. No project-wide dollar estimate is claimed.
Published downloads retain every model/tier/region group and every returned daily token aggregate. The web tables below show all model identifiers, monthly coverage, and service tiers. January and February 2026 have no returned token records—not confirmed zero spend.
Observed point coverage is 2025-10-30 00:00:57 → 2026-09-29 13:08:12 UTC (03:00:57 → 16:08:12 MSK). The active project was created in March 2023; the older empty queried windows are a lifetime-coverage gap.
Evidence was checked independently, not copied from a dashboard
Cloud retrieval completed 1,094 bounded/overlap queries and 1,316 pages. It removed 38,161 duplicate points with zero value conflicts. A separate nine-page unaligned whole-period query exactly matched the input/output totals. Parent verification re-summed all 682,296 unique raw token points and the 44 model/tier/region groups. A conflicting aligned/resampled diagnostic was excluded.
Local evidence covers the default database; there is no profiles directory on this host. A September backup added no new usage (80 exact duplicate rows); a June archive had no Vertex/Gemini metadata matches. All 384 confirmed rows in the earlier broad extract match. Its other three Gemini rows have an unknown route, explaining 387 versus 384.
Local storage has 90-day retention and usage rows cascade with session deletion. It is not a lifetime ledger. Two local rows spanning MSK dates represent 74 calls and remain unallocated by day. Other daily allocations use recording dates, not proved billing dates. Cloud daily/monthly tables use UTC point-end dates; their time basis is explicitly different.
Billing access is the remaining blocker to a true total
The active project's billing is enabled, but reading its linked billing account returns 403 PERMISSION_DENIED. No billing export tables were found in visible, fully paginated BigQuery datasets; hidden exports may exist. The browser requires a Google sign-in. Another accessible account's currency does not establish this project's invoice currency.
To close the question, an authorized billing viewer must provide the project's cost report or existing billing export, including gross charges, credits, invoice currency, and taxes; invoices/payment records are needed for the amount actually paid. Google distinguishes project cost-view permission (billing.resourceCosts.get) from billing-account spend-view permission (billing.accounts.getSpendingInformation).[3] No billing/IAM changes or new paid exports were made.
A bill can include cache storage, grounding, training, hosted capacity, audio/image SKUs, and unrelated services. Preserve these separately rather than treating token usage as the complete invoice. Reconcile the project bill against local attribution before calling it personal Hermes spend.
Download and reproduce without exposing conversation data
The report offers sanitized CSVs for all local rows, model/task summaries, recording-day allocations, all cloud model/tier/region groups, and daily/monthly cloud tokens; calculation JSON includes exact rates and assumptions. Raw API responses and scoped local evidence are preserved privately, with hashes. No prompts, generated text, credentials, project IDs, session IDs, or endpoint URLs are included in the public downloads.
No inference was invoked for the audit, and no runtime model configuration was changed. This is a static snapshot; no recurring monitor was enabled.
Sources
[1] https://cloud.google.com/vertex-ai/generative-ai/pricing — Official Vertex model pricing > "| Gemini 3.5 Flash | Input (text, image, video, audio) | Global | $1.50 | $1.50 | $0.15 | $0.15 |" > "Promotional pricing provided through 50% credits back on net spend on select models within a given period." > "| Gemini 3.8 Flash* through December 31, 2026 | Input (text, image, video, audio) | Global | $0.75 | $0.75 | $0.075 | $0.075 |" [2] https://docs.cloud.google.com/monitoring/api/metrics_gcp_a_b — Official Google Cloud Monitoring metric definitions > "Accumulated input/output token count." [3] https://docs.cloud.google.com/billing/docs/how-to/billing-access — Cloud Billing viewing permissions > "To grant permission to a user to view the costs of all projects under a Cloud Billing account, give the user permission to view the costs for a Cloud Billing account"