US지정학·Yahoo Finance RSS·

1.3조 달러 추론 전쟁 심화, AI 인프라 경쟁 가열

The $1.3 Trillion Inference War Is Heating Up. 3 Stocks to Watch.

2026.08.12 17:15 번역됨
AI 감성 분석
롱 (매수 신호)
롱 76%숏 24%

1.3조 달러 규모의 추론 시장에서 경쟁이 심화되면서 엔비디아와 AMD와 같은 기존 선두 주자들에게 강력한 상승 동력이 발생합니다.

핵심 요약

2032년까지 1.3조 달러 규모로 성장할 추론 시장에서 엔비디아, 세레브라스, AMD가 시장 점유를 위해 경쟁하고 있습니다.

핵심요약

  • AI 추론 시장은 2032년까지 1.3조 달러 규모로 성장할 것으로 전망됩니다.
  • 엔비디아는 Groq 및 LPU(언어 처리 장치) 인수를 통해 추론 시장에 진출했습니다.
  • 세레브라스는 SRAM 기반 칩을 사용하여 LPU 대비 5~6배 빠른 성능을 제공합니다.
  • 추론은 Raw Compute보다 빠른 메모리 접근과 낮은 지연 시간(latency)이 핵심입니다.

도입

본 기사는 AI 인프라 시장의 가장 빠르게 성장하는 부분인 추론(Inference) 시장의 잠재적 가치와 그 경쟁 구도를 조명합니다. 이 시장의 규모가 2032년까지 1.3조 달러에 달할 것으로 예측됨에 따라, 선두 기업들은 단순한 컴퓨팅 파워 경쟁을 넘어 추론 효율성을 극대화하는 새로운 아키텍처를 통해 시장 지배력을 확보하려 하고 있습니다. 투자자들은 이 경쟁이 향후 AI 인프라 시장의 주요 성장 동력이 될 것이라는 점을 주목해야 합니다.

본문 1: 추론 시장의 구조적 변화와 엔비디디아의 전략

추론 시장은 단순한 AI 모델 학습(Training)을 넘어 빠른 메모리 접근과 낮은 지연 시간(Latency)이 핵심 경쟁 요소인 영역입니다. 엔비디아는 이러한 추론 요구사항에 대응하기 위해 혁신적인 접근 방식을 취하고 있습니다. 엔비디아는 GPU의 그래픽 처리 장치(GPU)가 프롬프트 읽기(pre-fill) 단계를 처리하고, LPU(언어 처리 장치)가 응답 생성(decode) 단계를 처리하도록 시스템을 설계하여 응답 시간을 단축시킵니다. 이는 추론 과정의 각 단계에 최적화된 시스템을 제공함으로써, 학습 시장에서의 우위를 추론 시장으로 확장하려는 전략입니다. 이러한 접근 방식은 엔비디아가 추론 시장에서 선도적인 인프라 리더십을 유지하는 데 기여할 것으로 보입니다.

본문 2: 경쟁사들의 차별화된 아키텍처와 기술적 트레이드오프

엔비디아 외에도 경쟁사들은 추론 효율성을 높이기 위해 각기 다른 기술적 경로를 선택하고 있습니다. 세레브라스는 LPU와 유사하게 SRAM(정적 랜덤 액세스 메모리) 기반 칩을 활용하지만, 물리적 크기가 큰 SRAM의 제약을 극복하기 위해 웨이퍼 크기의 칩을 구현했습니다. 이로 인해 세레브라스 칩은 LPU보다 5배에서 6배 더 빠른 속도를 달성할 수 있습니다. 그러나 이러한 고성능을 구현하는 과정에서 세레브라스는 특화된 냉각 및 에너지 관리 솔루션이 필요하다는 기술적 트레이드오프를 감수해야 합니다. 반면, AMD와 같은 다른 칩 제조사들은 기존 컴퓨팅 아키텍처에 추론 최적화 기능을 통합하는 방식으로 경쟁하며, 이는 하드웨어 설계의 효율성과 시스템 통합 능력에 따라 차별화될 것입니다.

본문 3: 시장 경쟁의 장기적 전망과 리스크

이러한 추론 인프라 경쟁은 단기적인 성능 향상을 넘어 장기적인 시스템 통합 및 에너지 효율성 측면에서 중요한 의미를 가집니다. 추론 시장의 성장은 메모리 대역폭과 온칩(On-chip) 메모리 기술의 발전에 직접적으로 연관되어 있습니다. 향후 경쟁은 단순히 칩의 연산 능력 경쟁이 아니라, 메모리 계층 구조와 시스템 레벨의 통합 효율성을 얼마나 잘 달성하느냐에 달려 있습니다. 따라서 각 기업은 추론 특화 메모리 솔루션과 에너지 효율적인 칩 디자인에 대한 투자를 지속해야 합니다. 만약 특정 기업이 추론 특화 메모리 솔루션에서 독점적인 기술을 확보한다면, 이는 전체 AI 인프라 시장에서 강력한 해자(Moat)를 구축하는 기반이 될 것입니다. 다만, 세레브라스와 같이 대형 칩을 구현할 때 발생하는 냉각 및 전력 관리의 복잡성은 시장 진입의 장벽으로 작용할 수 있으므로, 기술적 우위와 상업적 실현 가능성 사이의 균형을 찾는 것이 중요합니다.

결론

결론적으로, 1.3조 달러 규모의 추론 시장은 메모리 접근 속도와 시스템 통합 효율성이 핵심 경쟁 요소가 되는 새로운 패러다임으로 진화하고 있습니다. 엔비디아의 시스템 통합 능력과 세레브라스의 대규모 SRAM 기술은 추론 시장의 미래를 형성할 잠재력을 가지고 있습니다. 향후 시장의 흐름은 특정 칩 제조사의 독점 기술 확보 여부와 시스템 레벨의 에너지 효율성 달성 능력에 따라 결정될 것으로 전망됩니다. 투자자들은 이러한 기술적 차별화와 시장 점유율 변화에 주목해야 합니다.


원문 링크: https://www.fool.com/investing/2026/08/12/trillion-inference-war-heating-up-stocks-nvda-amd/?.tsrc=rss

Original Article

The $1.3 Trillion Inference War Is Heating Up. 3 Stocks to Watch.

Inference has become the fastest-growing part of the artificial intelligence (AI) infrastructure market, and Bloomberg Intelligence projects it will double the size of the AI training market by 2032, reaching $1.3 trillion. With so much at stake, both leading chipmakers and upstarts are jockeying to grab a slice of this huge, fast-growing market.

Nvidia ( NVDA -0.02% ) , Cerebras ( CBRS +2.06% ) , and Advanced Micro Devices ( AMD +1.01% ) are all tackling this market in different ways. Let's see how these AI stocks stack up and why they could all be winners, given the size and growth of the inference market.

Already the winner in AI model training, Nvidia now has its sights on the inference market. The company's big move to capture share was its "acquisition" of Groq and its language processing units (LPUs). Inference is more about fast memory access and low latency than raw compute power, and LPUs help address this by having SRAM (static random-access memory) embedded directly on their chips.

LPUs are particularly useful during the decode phase of inference, which is when large language models (LLMs) answer queries. As such, Nvidia now offers complete systems designed specifically for inference, where its graphics processing units (GPUs) handle the pre-fill phase (reading the prompt) while its LPUs handle the decode phase, thereby speeding up response times.

This is a nice solution and positions Nvidia to remain an AI infrastructure leader, even if it doesn't capture the same market share it does in training.

Like Nvidia, Cerebras is tackling inference with SRAM-based chips. However, because SRAM is so bulky, instead of just embedding a small amount onto its chips and stringing them together, Cerebras has created huge wafer-sized chips that are five to six times faster than LPUs.

The physical size of Cerebras' chips comes with some trade-offs. They require specialized cooling and energy management solutions and, as such, are only sold or rented as part of the Cerebras CS-3 systems. They also come at a very premium price tag.

However, the company has inked major deals with OpenAI and Amazon 's AWS, and it recently announced a partnership with AMD that should help reduce the cost of ownership. The two companies will offer an inference solution in which AMD's Helios rack-scale solution will handle the pre-fill phase of inference, which it can do more cheaply, while Cerebras' Wafer-Scale Engine will perform the decode phase, which it can do more quickly. It's a nice way for companies to better compete with Nvidia's offerings.

Given the high cost of its systems, Cerebras has been more of a premium, niche solution, but it now looks set to become a major player in the humongous inference market.

After losing out on the LLM training market to Nvidia, AMD has been aggressively pursuing the inference market to make sure it doesn't get left behind again. Its chiplet design is better suited for inference, as it allows its GPUs to be packaged with more high-bandwidth memory (HBM) and to act as part of an entire unit to reduce latency. Meanwhile, its partnership with Cerebras looks like a smart move to help it better compete with Nvidia's complete inference system.

However, the company has not stopped there. It recently acquired memory optimization company MEXT and chip start-up Taalas to boost its inference offering. Memory is one of the biggest AI bottlenecks right now, and MEXT's solution can offload seldom-accessed data from DRAM to unused flash and then, using predictive AI, can transfer it back into DRAM before an application even requests it. This can reduce the need for more expensive DRAM and help save costs.

Meanwhile, Taalas has developed chips in which AI models are hardwired directly to bolster inference performance. Since the chips are model-specific, they aren't as flexible, but they are cheaper and much faster. The company plans to use them as part of a complete system where its GPUs would handle the pre-fill phase and Taalas' chips would handle the decode phase.

AMD is tackling inference from a couple of different angles, which should position it to capture a nice share of this huge market.

Source: https://www.fool.com/investing/2026/08/12/trillion-inference-war-heating-up-stocks-nvda-amd/?.tsrc=rss

주린이 © 2026