All Vertex usage / audit snapshot / 29 Sep 2026

The $0.75 was only compression.

This audit includes all retained Hermes tasks and every model in the accessible project's history. It does not turn usage counters into an invoice.

Actual total billed / paid
Unknown

Billing-account access returns 403.
No verified invoice or cost export.

Attributable Hermes usage
$108.71–$116.39

Two Standard-rate USD scenarios.
6,189 priced calls; 2 more unpriced.
6 Aug–29 Sep 2026.

Entire cloud project
21.75B

Recorded tokens, not dollars.
30 Oct 2025–29 Sep 2026.
Includes other applications.

No double counting. Hermes is a subset of the project inventory. The dollar figures are conditional scenarios, not invoice bounds. The project-wide bill and lifetime spend are still unverified.
Attributed local usage

All four models and all six tasks

Scenario A prices stored output only. Scenario B also adds separately recorded reasoning. Historical receipts do not prove one universal output-accounting rule.

ModelLocal callsA / USDB / USD
gemini-3.5-flash624$28.23$30.90
gemini-3.6-flash17$0.25$0.27
gemini-3.7-flash2$0.00$0.01
gemini-3.8-flash5,543$80.10$85.09

Local subtotal: $108.591623325 / $116.265996075. Add $0.12264675 for three temporally recovered calls: $108.714270075 / $116.388642825. Two missing receipts add an unknown amount, not zero.

Main chat carries almost all local cost

TaskLocal callsA / USDB / USD
main agent loop5,877$104.76$112.04
background review137$3.13$3.31
compression21$0.63$0.63
approval106$0.05$0.19
title generation44$0.02$0.09
vision1$0.01$0.01

The compression row has 21 local calls. Including its three recoveries gives the previous $0.75180375 compression estimate—already included in the combined amount.

Before advertised promotional credits: $189.20 / $201.88. These are alternative views of the same usage, not extra charges.

Assumes global Standard PayGo rates published on 29 September and eligible promotional credits. Local tier is not recorded; the project also has Priority use. Rates, credits, two missing calls, older unrecorded use, and taxes are not fully reconciled.

Project-wide cloud counters

All 20 model identifiers—not just Hermes

17,565,500,860 input + 4,185,586,581 output = 21,751,087,441 tokens. Nineteen model identifiers have nonzero tokens. Seven accessible projects were queried; one returned this inventory.

ModelInput tokensOutput tokens
gemini-3-flash-preview6,833,045,426268,018,488
gemini-2.5-flash4,761,769,5701,519,607,899
gemini-2.5-pro2,731,583,4722,285,126,005
gemini-3.1-flash-lite1,071,605,67124,779,396
gemini-3.5-flash653,399,59134,434,007
gemini-3.8-flash519,112,4208,178,999
gemini-3.1-pro-preview458,166,57418,099,517
gemini-2.5-flash-lite304,455,9756,305,448
gemini-3.5-flash-lite137,845,72912,673,833
gemini-flash-latest47,650,2953,206,311
gemini-3.7-flash29,953,0431,984,884
gemini-3.6-flash14,004,2183,149,260
gemini-embedding-0011,429,0100
gemini-live-2.5-flash-native-audio1,400,41813,549
gemini-2.0-flash76,9921,397
gemini-3.1-flash-image2,1775,577
gemini-2.5-flash-image571,290
gemini-flash-lite-latest215720
gemini-2.5-computer-use-preview-10-202571
gemini-omni-1.1-flash-preview00

20 identifiers do not imply 20 distinct underlying models: latest aliases and preview names are preserved exactly. Zero-token series remain visible. Project activity is not established as personal or Hermes-only usage.

Service tiers are not interchangeable
Tier labelInput tokensOutput tokens
shared unspecified188,747,14748,811,640
priority legacy label13,677,5357,068,554
standard14,957,028,7784,067,507,922
dedicated8,645,0321,120,026
priority2,397,402,36861,078,439
priority downgraded00

No Flex-labeled series returned. Unspecified is not Standard; dedicated capacity cannot be priced as PayGo. Missing cache, modality, context-size and thought detail prevents a defensible project-wide dollar total.

Project coverage reaches October 2025

Month / UTCInput tokensOutput tokens
2025-108,928,0722,467,578
2025-1125,034,5427,666,331
2025-1243,543,59716,202,719
2026-01No returned records; not confirmed zero
2026-02No returned records; not confirmed zero
2026-0383,679,35521,883,354
2026-04127,762,34818,912,949
2026-051,980,503,9362,144,309,224
2026-061,897,235,289456,426,613
2026-073,256,338,802474,738,486
2026-084,028,926,096464,320,252
2026-096,113,548,823578,659,075

Point-end timestamps, UTC. First point: 30 Oct 2025 00:00:57; last: 29 Sep 2026 13:08:12. Creation-to-cutoff queries extend into March 2023, but older empty periods do not prove zero spend.

1,982,469 invocations across all status codes; 1,971,056 HTTP 200. Invocation counts are not invoices.

A real total needs billing data, not another token guess

Project access works. Its linked billing-account read returns 403 PERMISSION_DENIED. No billing export was visible, and the browser requires sign-in. Actual currency, credits, tax, invoiced charges and payments remain unknown.

An authorized billing viewer must provide the project's cost report or existing export; payment/invoice records establish what was actually paid. No new permission, export, charge, or runtime change was made by this audit.

Cloud Billing has separate project-cost and billing-account spend-view permissions. Official access documentation ↗

The calculation is downloadable

Only sanitized usage metadata. No prompts, credentials, session IDs, project IDs or endpoint URLs. Missing values stay unknown; dates are labeled as cloud UTC or local recording-day MSK.

Full methodology, assumptions, verification, and sources

Two overlapping scopes answer different questions

Hermes-attributed usage: 6,186 calls in 384 retained usage rows, August 6–September 29, 2026. Three compression calls recovered from Monitoring bring priced coverage to 6,189 calls. Two further logged compression calls have no usable token receipt. Three other Gemini calls have no provider or endpoint and remain excluded from confirmed Vertex usage.

Entire accessible cloud project: 21,751,087,441 recorded tokens across 20 model identifiers (19 with nonzero tokens), October 30, 2025–September 29, 2026. This includes traffic from other applications; it is not proven to be Pavel's personal or Hermes-only consumption. It includes the local calls. Never add these two scopes.

All seven projects visible to the existing cloud credential were queried from creation to the cutoff. Only one returned usage. No returned data does not establish zero lifetime spend, and inaccessible projects remain outside coverage.

The usable Hermes cost estimate is conditional—not money proved paid

Using the published global Standard PayGo rates, correct cache buckets, and advertised promotional credits, the two output-accounting scenarios are US$108.71 and US$116.39. Exact values: $108.714270075 / $116.388642825. Both include $0.12264675 for the three temporal recoveries.[1]

  • Scenario A: price stored completion tokens only. This can omit thinking tokens.
  • Scenario B: price stored completion plus separately stored reasoning. This can double count if a receipt already included thoughts.
  • Before advertised credits: the same scenarios are US$189.20 / US$201.88. Do not add pre-credit and promotional figures together.
  • Actual billed and paid amount: unknown. These scenarios are neither an invoice nor lower/upper bounds. They omit unknown calls, unrecorded older usage, non-token charges, taxes, and any account-specific adjustments.
  • Standard tier is an explicit assumption for local receipts, which have no retained service-tier field. The broader project has Priority traffic. The audit uses published prices as of September 29, not verified historical invoice SKUs.

All four locally attributed models are included

Model Calls Scenario A Scenario B
gemini-3.5-flash 624 $28.23 $30.90
gemini-3.6-flash 17 $0.25 $0.27
gemini-3.7-flash 2 $0.00 $0.01
gemini-3.8-flash 5,543 $80.10 $85.09
Local subtotal 6,186 $108.59 $116.27
Additional temporal recoveries: 3.8 Flash 3 $0.12 $0.12
Combined priced coverage 6,189 $108.71 $116.39

Displayed lines round individually; exact CSV/JSON values are used for totals. The two unpriced calls are not included in these amounts.

Main chat—not compression—dominates local use

Task Calls with local usage Scenario A Scenario B
main agent loop 5,877 $104.76 $112.04
background review 137 $3.13 $3.31
compression 21 $0.63 $0.63
approval 106 $0.05 $0.19
title generation 44 $0.02 $0.09
vision 1 $0.01 $0.01

The compression row is the local 21-call subset. Its three temporal recoveries add $0.12264675, giving the earlier 24-call $0.75180375 estimate. Two logged calls remain unpriced. The $0.75 is already contained in the combined Hermes estimate; do not add it again.

Cache and reasoning accounting change the answer

Retained local counters: 77,707,943 uncached input, 424,526,358 cache-read, 0 cache-write, 1,139,081 stored output, and 1,631,146 separately recorded reasoning tokens. Source-code inspection proves local input is already net of cache. The reconstructed prompt volume is 502,234,301 tokens; subtracting cache again would undercount.

The historical Vertex OpenAI-compatible route stores completion tokens unchanged and reasoning separately. In 291 of 384 rows, reasoning exceeds output, including 67 single-call rows. Therefore a blanket claim that reasoning is already included is contradicted by these receipts. The native Gemini adapter uses a different route and cannot prove the historical Vertex accounting contract. The two scenarios expose that ambiguity instead of hiding it.

For each model, the calculation is (uncached input × input rate + cache reads × cache rate + stored output × output rate) / 1,000,000; Scenario B adds reasoning × output rate / 1,000,000. Cache-write is zero in these records.

Global Standard rates per million: Gemini 3.5 Flash $1.50 input / $0.15 cached / $9.00 output; 3.6, 3.7 and 3.8 Flash $0.75 / $0.075 / $3.75 during the promotion. Output pricing covers response and reasoning. Google describes the promotion as 50% credits back; actual application is unverified.[1]

The local database's historical estimated-cost sum, $92.830173375, is retained only as an unverified old estimate. It is not the newly calculated result or an invoice. Zero actual-cost fields without billing provenance were treated as unknown.

The project-wide inventory is much larger—and not personally attributable

The independently checked Monitoring inventory contains 17,565,500,860 input and 4,185,586,581 output tokens, plus 1,982,469 invocations across all response codes, of which 1,971,056 returned HTTP 200. Failed/status-diverse invocations are not all billable calls.

Monitoring's token counter is an accumulated input/output count, not a cost ledger.[2] The returned series lack cache quantities, modality, and per-request context tiers. They also include dedicated capacity, unspecified shared tiers, and legacy Priority labels. Pricing all input as uncached Standard or dedicated usage as PayGo would fabricate a project dollar total. No project-wide dollar estimate is claimed.

Published downloads retain every model/tier/region group and every returned daily token aggregate. The web tables below show all model identifiers, monthly coverage, and service tiers. January and February 2026 have no returned token records—not confirmed zero spend.

Observed point coverage is 2025-10-30 00:00:57 → 2026-09-29 13:08:12 UTC (03:00:57 → 16:08:12 MSK). The active project was created in March 2023; the older empty queried windows are a lifetime-coverage gap.

Evidence was checked independently, not copied from a dashboard

Cloud retrieval completed 1,094 bounded/overlap queries and 1,316 pages. It removed 38,161 duplicate points with zero value conflicts. A separate nine-page unaligned whole-period query exactly matched the input/output totals. Parent verification re-summed all 682,296 unique raw token points and the 44 model/tier/region groups. A conflicting aligned/resampled diagnostic was excluded.

Local evidence covers the default database; there is no profiles directory on this host. A September backup added no new usage (80 exact duplicate rows); a June archive had no Vertex/Gemini metadata matches. All 384 confirmed rows in the earlier broad extract match. Its other three Gemini rows have an unknown route, explaining 387 versus 384.

Local storage has 90-day retention and usage rows cascade with session deletion. It is not a lifetime ledger. Two local rows spanning MSK dates represent 74 calls and remain unallocated by day. Other daily allocations use recording dates, not proved billing dates. Cloud daily/monthly tables use UTC point-end dates; their time basis is explicitly different.

Billing access is the remaining blocker to a true total

The active project's billing is enabled, but reading its linked billing account returns 403 PERMISSION_DENIED. No billing export tables were found in visible, fully paginated BigQuery datasets; hidden exports may exist. The browser requires a Google sign-in. Another accessible account's currency does not establish this project's invoice currency.

To close the question, an authorized billing viewer must provide the project's cost report or existing billing export, including gross charges, credits, invoice currency, and taxes; invoices/payment records are needed for the amount actually paid. Google distinguishes project cost-view permission (billing.resourceCosts.get) from billing-account spend-view permission (billing.accounts.getSpendingInformation).[3] No billing/IAM changes or new paid exports were made.

A bill can include cache storage, grounding, training, hosted capacity, audio/image SKUs, and unrelated services. Preserve these separately rather than treating token usage as the complete invoice. Reconcile the project bill against local attribution before calling it personal Hermes spend.

Download and reproduce without exposing conversation data

The report offers sanitized CSVs for all local rows, model/task summaries, recording-day allocations, all cloud model/tier/region groups, and daily/monthly cloud tokens; calculation JSON includes exact rates and assumptions. Raw API responses and scoped local evidence are preserved privately, with hashes. No prompts, generated text, credentials, project IDs, session IDs, or endpoint URLs are included in the public downloads.

No inference was invoked for the audit, and no runtime model configuration was changed. This is a static snapshot; no recurring monitor was enabled.

Sources

[1] https://cloud.google.com/vertex-ai/generative-ai/pricing — Official Vertex model pricing > "| Gemini 3.5 Flash | Input (text, image, video, audio) | Global | $1.50 | $1.50 | $0.15 | $0.15 |" > "Promotional pricing provided through 50% credits back on net spend on select models within a given period." > "| Gemini 3.8 Flash* through December 31, 2026 | Input (text, image, video, audio) | Global | $0.75 | $0.75 | $0.075 | $0.075 |" [2] https://docs.cloud.google.com/monitoring/api/metrics_gcp_a_b — Official Google Cloud Monitoring metric definitions > "Accumulated input/output token count." [3] https://docs.cloud.google.com/billing/docs/how-to/billing-access — Cloud Billing viewing permissions > "To grant permission to a user to view the costs of all projects under a Cloud Billing account, give the user permission to view the costs for a Cloud Billing account"