Skip to main content
Xinying ZhisuanGPU SERVICE

Hard engineering behind your compute

Your GPU compute service partner

Modern compute is, fundamentally, GPU compute. Xinying Zhisuan works at the core of accelerator hardware — diagnostics, chip-level repair, genuine parts and performance tuning. Factory-grade craft, auditable process, and cards that go back into production.

BeijingShanghaiShenzhen7×24 fault response

years on GPU accelerators
20+years on GPU acceleratorsTeam drawn from tier-one manufacturers
primary platforms serviced
H100 / H200primary platforms servicedIncluding baseboards and NVSwitch
GPU recovery rate
95%GPU recovery rateBacked by an ISO cleanroom
response, three cities
7×24response, three citiesBeijing · Shanghai · Shenzhen

What we do

From a single board to an entire AI datacenter

Four service lines covering the full lifecycle of compute hardware — diagnose, repair, verify, tune, and supply.

  • Hardware diagnostics

    Hardware health screening covers core component triage and circuit integrity verification. On the software side we run BIOS, driver and management-tool compatibility checks, ECC validation, log monitoring and compliance auditing. Every engagement ends in a full-chain report.

    • ECC
    • BIOS
    • Compliance audit
    • Full report
  • Chip-level repair

    End-to-end repair from passive components to high-end silicon: GPU die reflow and reballing, VRAM package diagnosis and replacement, PCB microcrack detection and repair, thermal and liquid-cooling maintenance, gold-finger replating, high-speed trace repair, and vBIOS reflashing.

    • BGA reball
    • VRAM swap
    • PCB microcrack
    • Liquid cooling
    • vBIOS
  • Genuine parts supply

    100% original parts, sourced exclusively through authorised distributors, with complete purchase contracts, supply certificates and invoices retained for traceability. No refurbished pulls, no grey-market assemblies. Warehousing in Beijing, Shanghai and Shenzhen with 7×24 logistics.

    • Traceable
    • Three warehouses
    • 7×24 logistics
  • Reconfiguration & tuning

    A closed loop from safe teardown and precision replacement through custom upgrades to acceptance testing: memory expansion, PCIe lane bandwidth bottleneck analysis, GPU driver and CUDA environment tuning, deep-learning framework configuration, and OS kernel parameter adjustment.

    • CUDA tuning
    • PCIe bandwidth
    • Kernel params
    • MFU gains

Technical direction

AI fault localisation: turning repair experience into a reusable model

GPU fault localisation has always depended on senior engineers. The more experienced the engineer, the more accurate the call — but expertise does not clone, does not parallelise, and does get tired. Our direction is to make that reasoning chain explicit and learnable, so machines can carry the first layer of judgement.

The following describes Xinying Zhisuan’s R&D direction and methodology — the capability path we are building. It is not a commercially available product and constitutes no delivery commitment.

  • A real fault dataset

    Our repair system forms a closed data loop from work-order creation to final delivery: incoming inspection, initial test, quotation, repair, final test, acceptance — each stage producing structured records and measurement data. What accumulates is labelled fault samples with confirmed remedies and re-test verification. That is not something you can buy.

    • Structured orders
    • Labelled loop
    • Verified outcomes
  • Multimodal signal capture

    Fault signatures do not live only in logs. We capture software telemetry (DCGM temperature, power and ECC correction rates; Fieldiag diagnostic codes; GPU_burn sustained-load curves) alongside physical evidence (infrared thermal fields, X-ray solder imaging, oscilloscope waveforms) — a multimodal feature set for the same card.

    • DCGM
    • Fieldiag
    • GPU_burn
    • Thermal imaging
    • X-ray
  • Formalising signal-tracing diagnosis

    Our CTO’s “signal-tracing diagnostic method” was adopted into Huawei’s global repair manual. At its core it is an inference path running backwards from an observed anomaly to the failing component. Our work is to express that expert reasoning as trainable labelling rules and decision structure — so the model learns method, not just correlation.

    • Expert rules
    • Causal paths
    • Interpretability
  • Local and private deployment

    Cluster logs and fault data are highly sensitive customer assets. The models are designed for on-premise and private deployment, with inference running inside the customer environment and data never leaving the site — consistent with our existing isolation practices, ISO27001 framework, and China’s Data Security Law and PIPL.

    • Data stays on-site
    • Offline inference
    • ISO27001
  1. Shorter time to localise

    Move the senior-engineer judgement forward to intake, compressing the initial-test-to-quote cycle.

  2. Higher first-time fix rate

    Fewer teardowns driven by misdiagnosis, and less secondary damage risk to precision components.

  3. Toward predictive maintenance

    Identify degradation signatures from long-run telemetry trends and warn before an outage.

Stack direction
  • Local LLM / small models
  • Time-series anomaly detection
  • Multimodal fusion
  • Expert-rule distillation
  • RAG
  • NVIDIA DCGM
  • CUDA

Capabilities

Why the hard cards end up here

Equipment, cleanroom, process and people — the recovery rate only holds when all four are in place.

  • Factory-grade engineering team

    Every member of the technical team comes from a tier-one manufacturer, with more than twenty years on GPU accelerators and 15+ years and ten thousand cards of repair experience. Work is tiered: senior experts own the repair framework and complex localisation; principal engineers, fluent in NVIDIA and AMD schematics, handle 90% of hardware diagnosis and repair.

    • Low: 1 day
    • Medium: 3 days
    • High: 7 days
  • ISO cleanroom facility

    Our repair floor is built to ISO cleanroom standard with ±0.1℃ / ±1%RH environmental control and modular clean cabins. Removing dust, static and humidity swings prevents secondary damage to precision components — lifting GPU recovery to 95%, more than 20% above the industry norm.

    • ±0.1℃
    • ±1%RH
    • Modular clean cabins
  • Three-stage QA

    Initial inspection → repair → 72-hour burn-in, with 100% genuine replacement parts. Main and sub-processes run in parallel, each stage emitting its own document: incoming inspection report, initial test report, quotation, repair report, final test report, and outbound inspection report.

    • 72h burn-in
    • 100% genuine parts
    • Six reports
  • Data security & isolation

    Large accounts get a dedicated enclosed repair zone, a ring-fenced technical team, fully traceable repair video, and confidentiality auditing to guarantee zero data leakage. Where required we build a dedicated repair workshop at a customer-designated site — physical isolation through to data control.

    • ISO27001
    • Dedicated workshop
    • Traceable video

Diagnostic & repair equipment

  • GPU single-card test bench
  • Automated BGA rework station
  • X-ray solder inspection
  • Infrared thermal imager
  • Microscopy workstation
  • Digital oscilloscope
  • Ultrasonic cleaning system
  • Constant-temperature preheat station
  • Forced-air drying oven

Process

Five stages, fully visible

From work-order creation to final delivery, a closed data loop — repair progress is visible in real time.

  1. 01

    Consultation

    Online intake, on-site unboxing inspection

  2. 02

    Localisation & quote

    Pinpoint the fault, issue quotation and contract

  3. 03

    Customer approval

    Assign the repair team, open the service ledger

  4. 04

    Stress testing

    Real datacenter load testing with a test report

  5. 05

    Return shipment

    Insured express return, three-month warranty

Service commitments

to respond
2 hoursto respondDedicated support team
standard repairs
48 hoursstandard repairsRepaired and shipped back
business days
10–15business daysComplex cases
warranty
3 monthswarrantyStandardised repair process

Hardware maintenance tiers

Chosen by fleet size and response requirement — from standardised support to a dedicated on-site team.

  • Standard

    Remote · 7×24×10

    For smaller fleets (under ~100 servers) or customers with moderate response requirements.

    Response
    Fault calls accepted 24 hours a day, seven days a week; an engineer is dispatched on site within 10 hours of response.
    Staffing
    No on-site personnel
    • On-site fault diagnosis and parts replacement
    • Repair completed and returned within 14 days of receipt
  • Pro

    On-site · 7×24

    For medium-to-large clusters (recommended for fleets of 100+ servers), with full-time on-site coverage.

    Response
    On-site engineer reaches the fault location within 1 hour of the request.
    Staffing
    Four on-site staff per 300 servers, at least one with 3+ years of AI-hardware repair experience
    • Repair completed and returned within 7 days of receipt
    • Weekly fleet health inspection report with early risk warnings
  • MAX

    For strategic accounts

    A bespoke tier for customers whose core business depends on AI compute — large internet firms and research institutions.

    Response
    On site with replacement parts within 1 hour.
    Staffing
    A permanent team of no fewer than two on-site engineers
    • A ≥500 m² repair base established within 10 km of the customer
    • Fault-unit turnaround compressed to 5 days
    • Daily fleet analysis reports with real-time risk warnings
    • Quarterly hardware performance optimisation recommendations

Case studies

Real faults, real remediation

Two representative engagements — one a physical-layer failure, one a thermal and power issue surfacing as software errors.

An enterprise customer, Beijing

H200 module baseboard failure

Reported symptomThe customer reported an H200 module failing to operate. Initial consultation pointed to a suspected baseboard fault, and the unit was shipped to Xinying for professional inspection.

Findings

  • Cold solder joints → signal transmission interrupted, progressive shutdown
  • Circuit short → current overload, burnt components, total failure
  • Root cause: sustained high-load operation without maintenance intervals

Actions taken

  • Fault acknowledged within 30 minutes, contingency plan activated
  • Senior engineer confirmed the suspected fault point with the customer by phone within 2 hours
  • Cold-joint repair: microscope localisation → constant-temperature precision resoldering → verification
  • Short-circuit repair: tracer path isolation → damaged trace removal → jumper wiring → low-voltage validation

Outcome

  • Full-dimension PCB inspection with professional equipment; validation passed
  • Module restored to service, stability 190%+
  • Averted 50,000+ RMB per day in losses; business continuity preserved

A datacenter operator, Xinjiang

H100 software error fault repair

Reported symptomThe customer reported software errors on an H100 accelerator affecting production systems. The team opened a remote triage and repair workflow.

Findings

  • Heavily dust-clogged cooling fan → reduced thermal efficiency, core temperature triggering protection and surfacing as software errors
  • Aged power-circuit components → unstable supply, indirectly causing abnormal software behaviour

Actions taken

  • Responded within 30 minutes; engineers narrowed the fault point with the customer
  • Thermal remediation: professional dust removal from the fan assembly, aged thermal paste replaced
  • Power circuit repair: aged components located with a circuit tester, replaced like-for-like and parameters re-verified
  • Software layer: GPU driver updated and related system settings optimised

Outcome

  • Validation passed with no software errors
  • GPU operating temperature normal, power delivery stable
  • Production systems restored to smooth operation

Industries served

Where our work runs

Our customers operate where compute continuity matters most — where a day of downtime costs far more than the repair.

  • Major internet companies
  • State-owned and joint-stock banks
  • The three telecom carriers
  • Server and system OEMs
  • AI datacenter and IDC operators
  • Research institutes and university labs

More than 30 well-known enterprises served to date, with 97% customer satisfaction.

Contact

Cards to repair, or a conversation about AI diagnostics

Describe the symptom and the card model and we will assign an engineer at the matching tier. Difficult cases can go straight to the technical team.

Business enquiries

15820792879

Mr. Dai

  • Corporate email

    Coming soon

  • Service hours

    7×24 fault response

    Service centres: Beijing · Shanghai · Shenzhen

  • Address
    Room A710, 136 Nanxin District, Qiaotou Community, Fuhai Street, Bao'an District, Shenzhen, Guangdong, China