
Identifying Opportunities to Dramatically Improve AI Efficiency
A UIDP Workshop
First virtual meeting: November 20, 9 am-1 pm PST
Second virtual meeting: December 4, 9 am-1 pm PST
In-person workshop: San Francisco, CA, December 9-11 (three full days)
San Francisco, CA
UIDP, with support from the U.S. National Science Foundation (NSF Award #2534301), will convene researchers, entrepreneurs, and commercial providers of AI to accelerate the adoption of translation-ready research results that can improve the overall efficiency of commercial AI/ML workloads. The premise of the workshop is that researchers have developed technologies that have the potential to collectively deliver multiple orders of magnitude in improvement in the efficiency with which AI/ML workloads are fielded. Rapid adoption of these technologies can, in turn, improve U.S. competitiveness in the AI space, e.g., by addressing near term limitations on energy supplies, accelerating time to market of new models, etc.
Key dates. location, and self-nomination form:
- First virtual meeting: November 20, 9 am-1 pm PST
- Second virtual meeting: December 4, 9 am-1 pm PST
- In-person workshop: San Francisco, CA, December 9-11 (three full days)
Application/nomination timeline: Self-nominations will be evaluated on a rolling basis with timely decision notifications until the limit of 35 participants balanced across technical areas is reached. For fullest consideration, participants are strongly encouraged to self-nominate by Friday, October 24. See here for the self-nomination form.
Format and Participant Profiles
The workshop will be organized to facilitate proposal teaming between: (a) participants with translation-ready results at various layers of the stack who are passionate about seeing their technology put to use; and (b) problem owners who come from organizations that deploy AI/ML systems at scale and seek to dramatically improve the efficiency and competitiveness of their offerings. In baseball jargon, we think of the first type of participants as pitchers and the second type, i.e., the problem owners, as catchers. Particularly sought as pitchers are entrepreneurial researchers (in academia, government, ventures, corporate research and/or non-profits) seeking to implement their research results at-scale in real-world deployments.
The outcome, by the final day of the workshop, will be a set of 5-10 concrete project outlines each championed by a team of participants that will have created a convincing presentation and a written project plan for implementing the research results at commercial scale. Examples of the target metrics are watt-hour and peak watt improvements for major AI training and inference workloads subject to latency, throughput, and/or other relevant constraints.
Technical Scope
The opportunity design space includes, but is not limited to, the right-sizing of the resources deployed (compute/memory/communications), algorithmic tradeoffs, and the customization of well-known systems techniques to at-scale AI/ML environments. Some of these techniques can be embraced at multiple layers of the stack and across the cloud-edge to significantly improve efficiency and reduce redundancy, suboptimality, and excess precision. There may be additional opportunities for efficiency gains through holistic system thinking, i.e., the adoption of a full-stack perspective.
Figure 1: Scope of component and/or system-level results for energy efficiency
As illustrated in the diagram, translation-ready results are specifically sought in the context of the software supporting AI/ML pipelines, its supporting hardware, and the mediating MLOps and distributed system software.
To elaborate on the terms in boldface:
- Translation-ready results: This workshop is a translational activity targeting dramatic near-term improvements in AI/ML efficiency. In-scope are only those concepts that do not require significant additional research and that do not depend on far-off standards, buildouts, or roadmap advances in slow-moving areas such as energy generation, batteries, power grid management, materials research, etc.
- Software implementation of AI/ML: In-scope ideas include software and algorithms that lead to improvements across the full AI/ML pipeline from data preparation to training to inference-serving to application architecture. These techniques could focus on the efficient utilization of computational resources and/or of memory/storage resources. Importantly, the goal is not to change the internals of the models (or sub-models) themselves but rather to improve the overall efficiency of their implementation and deployment. Automated techniques that involve mixtures of experts, the automated distillation of models, cloud-edge partitioning, etc., are within scope.
- Hardware supporting AI/ML: In-scope are improvements in the efficiency of the compute triad (logic, memory/storage, and communications) at any scale (processor, rack, datacenter, cloud/edge). Demand response flexibility, power distribution, and thermal management are also in scope. A specific challenge in the hardware space will be identifying paths through which promising innovations can rapidly be adopted.
- MLOps and distributed system software: Also in scope are improvements and new paradigms for system software (scheduling, orchestration) that can exploit improvements in the underlying infrastructure for joint optimization for a given workload. Importantly, this category is more concerned with improving interactions between nodes than within nodes.
Workshop Participant Roles and Registration of Interest Form
“Catchers”
UIDP welcomes inquiries from interested “catchers” of relevant technical results. The ideal participant profile is someone with responsibility for developing and/or serving frontier models at scale in an AI company or cloud infrastructure company. Applicants will typically be affiliated with commercial organizations.
“Pitchers”
UIDP welcomes inquiries from interested “pitchers” of relevant technical ideas. The ideal participant profile is a researcher or entrepreneur with translation-ready research results and a passion to see those results put into practice. Applicants may be affiliated with academia, non-profits, startups, and/or industry research labs.
When
- First virtual meeting: November 20, 9 am-1 pm PST
- Second virtual meeting: December 4, 9 am-1 pm PST
- In-person workshop: San Francisco, CA, December 9-11 (three full days)
Where
San Francisco, CA
