Find Jobs
Find Jobs Near You – Available Work in Your Location
Skip to job details
QI
quadric, Inc
AI Performance Modeling Engineer
Career Insights for Artificial Intelligence Engineer (General)
See where this job fits in the broader career landscape. Knowing your career path helps you see what's possible from here.
Scorecard
Based on California data
Review key factors to help you decide if this role fits your goals. How is this calculated?
What they do
An Artificial Intelligence Engineer develops, tests, and deploys artificial intelligence models. May work closely with data software engineers and data professionals to train and implement AI models into existing systems or develop new applications.
$160,491 / year median in California
Job Description
AI Performance Modeling Engineer quadric, Inc - 4.0 Burlingame, CA Job Details $150,000 - $200,000 a year 1 day ago Benefits Paid parental leave Health savings account Disability insurance Health insurance Dental insurance 401(k) Flexible spending account Paid time off Parental leave Vision insurance Life insurance Qualifications Content creation for technical audiences Performance optimization Technical documentation System architecture Computational modeling Computer hardware Quantitative analysis Technical writing experience within technology Python Full Job Description About Quadric Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C++ code on a single programmable architecture. Our technology powers intelligent edge devices across automotive, industrial, robotics, and embedded systems. Founded by technologists from MIT and Carnegie Mellon, Quadric is a well-funded growth-stage semiconductor IP company with a growing licensing business. As we enter our next phase of growth, we're looking for our first true marketing leader to build and scale the function. The Opportunity Quadric has created an innovative General-Purpose Neural Processing Unit (GPNPU) architecture. Unlike standard accelerators, the Quadric GPNPU executes both neural network graph code and conventional C++ DSP/control code across edge and endpoint devices. As an AI Performance Modeling Engineer , you will build analytical, cycle-level performance models of AI inference workloads on our next-generation architecture in Python before silicon exists. These models directly guide team decisions on hardware lane bindings, tensor placement, and architecture trade-offs. We welcome candidates across all experience levels—from early-career engineers to seasoned experts—with direct mentorship provided to help you master mapping complex workloads (like LLMs) onto our custom hardware. What You'll DoPerformance Modeling & Architectural Analysis Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware. Derive from first principles which hardware lanes operations bind on (compute, on-chip/external memory bandwidth, interconnect) and model software pipelining overlaps. Model tensor placement, tiling across processing elements, local memory residency, and data movement across memory tiers. Model sharding and collective boundary communication across multi-die systems. Workload Adaptation & Technical Writing Incorporate architectural details across vision networks and Large Language Models (LLMs), including operator mix, sparsity, routing, and quantization/low-precision numeric formats. Calibrate performance models against an instruction-set simulator and profiling traces to meet stated accuracy targets. Write and defend technical studies presenting empirical evidence that directly informs architecture and product decisions. Balance single-stream latency against scaled throughput performance. What Success Looks Like Within your first 6-12 months, you'll: Own a full workload's model end to end, calibrated against simulation and trusted by the engineering team. Build performance models that consistently predict workload behavior within 10-15% of actual measurements. Publish a written study whose defended conclusions directly shape an architecture or product decision. Review and extend performance models beyond your initial starting domain.