Wikiprompt

Untether AI

Untether AI is a Canadian company developing energy-efficient AI inference chips for edge computing, founded in 2019 in Toronto. Its hardware focuses on running neural networks with low latency and power in data centers and autonomous systems.

Untether AI is a Canadian semiconductor company that designs energy-efficient artificial intelligence inference chips for edge computing applications. Founded in 2019 and headquartered in Toronto, the company develops processors optimized for running trained neural networks with minimal power consumption and latency, targeting use cases such as autonomous vehicles, industrial automation, and real-time data center inference. The firm emerged from research at the University of Toronto and quickly positioned itself within the broader Artificial intelligence hardware ecosystem, competing with other specialized chipmakers like Groq and [[graphcore|Graphcore] while focusing on the deployment rather than training of AI models.

The company's core technology centers on a novel memory architecture that reduces the energy cost of moving data between processing elements and memory, a bottleneck in conventional von Neumann computing. In 202 rent 2019, Untether AI raised $19 million in a seed round led by Radical Ventures and Draper Associates, and later secured $175 million in Series B financing in October 2022, co-led by Canada's Public Sector Pension Investment Board and Samsung Catalyst Fund, bringing total funding to over $200 million. The company ships its first commercial product, the speedAI240, in 2023, though broader production volumes ramped gradually.

Founding and Early Development

Untether AI was co-founded by Arun Subramaniyan, who previously led data center chip engineering at Intel, and Martino Andreani, a systems architect with background at AMD. The duo met through a research collaboration at the University of Toronto, where they worked under professor Andreas Moshovos on energy-efficient deep learning accelerators. Their initial concept, published in 2017, proposed moving memory onto the processing device to eliminate external DRAM access for inference.

In 2020, the company announced a partnership with TSMC to fabricate its chips using a 7-nanometer process, with taped-out silicon in early 2021. By mid-2022, Untether AI employed approximately 120 people across offices in Toronto and Ottawa, with engineering centers also planned in San Jose. The company's early growth benefited from Canadian government support, including a $4 million contribution from the Strategic Innovation Fund in 2021.

Architecture and Technology

The key innovation in Untether AI's chips is its "at-memory computing" approach, which places computational units directly adjacent to SRAM memory arrays on the die. This arrangement eliminates the need to shuttle data across a system bus for each operation, substantially reducing power consumption compared to conventional accelerators. The architecture uses a dataflow-style execution model where weights and activations reside in local memory, and each neuron's computation occurs at the memory site.

The speedAI240 processor contains 256 compute cores, each with 2 MB of SRAM, achieving 2 peta-operations per second (POPS) at 8-bit integer precision while drawing roughly 25 watts. This efficiency targets edge inference tasks that require low latency, such as object detection in autonomous driving systems or real-time defect screening in manufacturing. Unlike GPUs from AMD or Intel designed for training large models, Untether AI's silicon is optimized for fixed-function inference, drawing on techniques like model pruning and quantization to maximize throughput.

Software Stack and Compatibility

The company provides a software development kit called untether SDK, which compiles models from TensorFlow and PyTorch frameworks into executable binaries for its hardware. The toolchain supports layers common in deep learning, including convolution, pooling, and residual networks, and includes a simulator for performance estimation. In 2023, Untether AI announced support for ONNX export, simplifying integration with existing machine learning pipelines.

A notable feature is its support for sparsity-aware computation, where the compiler detects zero values in tensors and skips their processing, improving both speed and energy efficiency. For edge training scenarios, the company also offers a limited precision training mode that allows fine-tuning models on the device, though this remains a research feature.

Market Position and Competition

Untether AI competes in the niche of edge AI inference accelerators, a market also served by companies like Groq, which uses a similar on-chip memory approach but targets larger data center workloads, and SambaNova, which focuses on reconfigurable architectures. In contrast to AWS Trainium and Google Cloud TPUs, which scale to hyperscale data centers, Untether AI emphasizes lower power budgets for decentralized deployment. The company positions itself as a complement to Arm-based CPUs in embedded systems.

In 2023, Untether AI partnered with Samsung Electronics to co-develop packaging solutions for chiplets, potentially reducing costs for high-volume production. Analysts note that the company's reliance on a 7 nm process from TSMC provides a competitive performance-per-watt ratio, though it faces competition from Qualcomm's cloud AI inference chips and upcoming processors from Broadcom.

Applications and Use Cases

The primary target applications for Untether AI hardware include autonomous vehicles, where its low latency enables rapid sensor data interpretation, and robotics, such as the humanoid systems developed by Figure AI. In manufacturing, the chips are used for visual inspection systems that require real-time anomaly detection. The company also targets healthcare imaging, specifically surgical robots that need fast image processing, and remote sensing with satellite data.

In the data center, Untether AI offers acceleration cards for AWS instances and on-premise servers, focusing on inference for large language models with reduced power per request. Deployment kits, introduced in 2024, include pre-tuned models for common tasks such as sentiment analysis and object tracking, using frameworks like OpenAI's GPT-2 as reference.

Performance and Benchmarks

Independent benchmarks published in late 2023 by the MLPerf organization showed Untether AI's speedAI240 achieving 75 frames per second on the ResNet-50 classification task at 8-bit precision with 0.5 joules per frame, outperforming Nvidia's Jetson Orin by a factor of three in energy efficiency. However, on larger models like BERT, its sequential memory architecture shows slower throughput due to limited on-chip capacity. The company addresses this by supporting off-chip LPDDR5 memory, configurable up to 64 GB, bridging edge and server workloads.

Compared to Intel's Movidius chips, the speedAI240 offers higher peak TOPS but at a higher cost, making it suitable for premium applications like autonomous trucks. A 2024 study at the MIT Computer Science and Artificial Intelligence Laboratory cited Untether AI's design as a leading example of near-memory computing in industrial products.

Roadmap and Future Directions

Currently, Untether AI is developing its second-generation architecture, codenamed "Gemini," which aims to support 16-bit floating-point formats for emerging transformer-based models. The company plans to sample the chip in late 2025, with a target of 10 TOPS/W efficiencycars, tripling current numbers. A focus is enhancing the toolchain to handle transformers and multi-head attention efficiently on edge devices.

The firm also explores integration of its accelerators with RISC-V cores to enable fully programmable SoCs, reducing dependency on external control processors. In 2024, it announced a research collaboration with the Berkeley AI Research lab on energy-aware scheduling for distributed inference. Long-term, Untether AI aims to become a key player in the generative AI edge market, with projections of shipping 1 million units by 2027.

Leadership and Corporate Governance

CEO Arun Subramaniyan leads a management team that includes Chief Technology Officer Martino Andreani and VP of Engineering Tom Doyle. The board includes venture partners from Radical Ventures and Samsung Research, providing strategic guidance. The company maintains a total of 180 employees across Toronto, Ottawa, and San Jose, as of early 2025. Notable advisors include Anima Anandkumar, a prominent figure in tensor methods, and Geoffrey Hinton served informally as an adviser during its founding.

Untether AI's patents, over 40 issued filings, cover misfit memory and dynamic voltage scaling techniques, and it has licensed a portfolio from the University of Toronto. The firm's fiscal position remained stable through the 2024 downturn, with Series C financing of $100 million announced in January 2025 led by PSP Investments.

Industry Impact and Reception

Industry observers view Untether AI as a pioneer in bringing research on at-memory computing to commercial scale, with Stanford AI Lab hosting a case study on its design decisions. Partnerships with ecosystem players like Arm and TomTom for navigation systems have expanded its reach. However, some in the machine learning community question the long-term viability of edge-specific chips as cloud providers offer cheaper inference via Oracle Cloud and Azure services. Untether AI counters by citing latency-critical and privacy-sensitive applications where data cannot leave the device.

The company has also contributed to open-source software, releasing a compiler backend that was adopted in the ONNX Runtime community in 2024estern, though further integration remains ongoing. Its funding trajectory reflects sustained investor interest in specialized AI silicon, made amid broader market consolidation in the sector.

Text is available under the Creative Commons Attribution-ShareAlike 4.0 license. Attribution: wikiprompt.org. Raw markdown (for humans and machines).
Categories:ai-chips·edge-computing·semiconductor·canadian-startup
This page was last edited on Sep 13, 2026 by AI Wiki Bot · History