Deploying Transformers at Scale

Feb 16, 2022

Share this post

Sharing to X

Sharing to Email

Key takeaways:

ONNXRuntime is the best inference package for Transformer networks;
Nvidia Triton, together with ONNXRuntime is the best solution for GPU inference;
Optimization matters. It's quite easy to unlock a >10X performance gain in 2022.

About this session

Transformer networks have taken the NLP world by storm, powering everything from sentiment analysis to chatbots. However, the sheer size of these networks presents new challenges for deployment, such as how to provide acceptable latency and unit economics.

The de-identification tasks Private AI services rely heavily on Transformer networks and involve processing large amounts of data. In this talk, I will go over the challenges we faced and how we managed to improve the latency and throughput of our Transformer networks, allowing our system to process Terabytes of data easily and cost-effectively.

Watch the full session:

https://www.youtube.com/watch?v=BgP6fEIadtgThis talk was originally delivered at the 2021 Toronto Machine Learning Summit.

Speaker bio: Pieter Luitjens is the Co-founder & CTO of Private AI. He worked on software for Mercedes-Benz and developed the first deep learning algorithms for traffic sign recognition deployed in cars made by one of the most prestigious car manufacturers in the world. He has over 10 years of engineering experience, with code deployed in multi-billion dollar industrial projects. Pieter specializes in ML edge deployment & model optimization for resource-constrained environments.

Contact us to request Pieter as a guest speaker.

Sign up for our Community API

The “get to know us” plan. Our full product, but limited to 75 API calls per day and hosted by us.

Get Started Today

Contact Us To Transform Your Data

Building Privacy into AI: The Strategic Case for Advanced PII Detection

Oct 14, 2025

Contact Centers, Chat, and Email: PII Detection in Customer Communications

Oct 9, 2025

Healthcare and Medical Data: The Ultimate PII Detection Challenge

Oct 7, 2025

The Specialization Gap: Purpose-Built vs. General Market PII Detection Solutions (Benchmark Results)

Oct 3, 2025

How to Properly Benchmark PII Detection Solutions: A Research-Based Methodology

Sep 22, 2025

The Hidden PII Detection Crisis: Why Traditional Methods Are Failing Your Business

Sep 16, 2025

Data Left Behind: AI Scribes’ Promises in Healthcare

Jun 18, 2025

Data Left Behind: Healthcare’s Untapped Goldmine

Jun 3, 2025

The Future of Health Data: How New Tech is Changing the Game

May 13, 2025

Why is linguistics essential when dealing with healthcare data?

May 8, 2025

Why Health Data Strategies Fail Before They Start

Apr 28, 2025

Private AI to Redefine Enterprise Data Privacy and Compliance with NVIDIA

Feb 20, 2025

EDPB’s Pseudonymization Guideline and the Challenge of Unstructured Data

Feb 14, 2025

HHS’ proposed HIPAA Amendment to Strengthen Cybersecurity in Healthcare and how Private AI can Support Compliance

Feb 13, 2025

Japan's Health Data Anonymization Act: Enabling Large-Scale Health Research

Feb 12, 2025

What the International AI Safety Report 2025 has to say about Privacy Risks from General Purpose AI

Feb 11, 2025

Private AI 4.0: Your Data’s Potential, Protected and Unlocked

Jan 27, 2025

How Private AI Facilitates GDPR Compliance for AI Models: Insights from the EDPB's Latest Opinion

Dec 20, 2024

Navigating the New Frontier of Data Privacy: Protecting Confidential Company Information in the Age of AI

Nov 14, 2024

Belgium’s Data Protection Authority on the Interplay of the EU AI Act and the GDPR

Nov 13, 2024

Enhancing Compliance with US Privacy Regulations for the Insurance Industry Using Private AI

Nov 13, 2024

Navigating Compliance with Quebec’s Act Respecting Health and Social Services Information Through Private AI’s De-identification Technology

Oct 29, 2024

Unlocking New Levels of Accuracy in Privacy-Preserving AI with Co-Reference Resolution

Oct 25, 2024

Strengthened Data Protection Enforcement on the Horizon in Japan

Oct 10, 2024

How Private AI Can Help to Comply with Thailand's PDPA

Oct 9, 2024

How Private AI Can Help Financial Institutions Comply with OSFI Guidelines

Sep 24, 2024

The American Privacy Rights Act – The Next Generation of Privacy Laws

Sep 19, 2024

How Private AI Can Help with Compliance under China’s Personal Information Protection Law (PIPL)

Sep 17, 2024

PII Redaction for Reviews Data: Ensuring Privacy Compliance when Using Review APIs

Sep 13, 2024

Independent Review Certifies Private AI’s PII Identification Model as Secure and Reliable

Sep 10, 2024

To Use or Not to Use AI: A Delicate Balance Between Productivity and Privacy

Aug 29, 2024

News from NIST: Dioptra, AI Risk Management Framework (AI RMF) Generative AI Profile, and How PII Identification and Redaction can Support Suggested Best Practices

Aug 20, 2024

Handling Personal Information by Financial Institutions in Japan – The Strict Requirements of the FSA Guidelines

Jul 12, 2024

日本における金融機関の個人情報の取り扱い - 金融庁ガイドラインの要件

Jul 12, 2024

Leveraging Private AI to Meet the EDPB’s AI Audit Checklist for GDPR-Compliant AI Systems

Jun 27, 2024

Who is Responsible for Protecting PII?

Jun 27, 2024

How Private AI can help the Public Sector to Comply with the Strengthening Cyber Security and Building Trust in the Public Sector Act, 2024

Jun 17, 2024

A Comparison of the Approaches to Generative AI in Japan and China

Jun 14, 2024

Updated OECD AI Principles to keep up with novel and increased risks from general purpose and generative AI

Jun 12, 2024

Is Consent Required for Processing Personal Data via LLMs?

Jun 10, 2024

The evolving landscape of data privacy legislation in healthcare in Germany

Jun 7, 2024

The CIO’s and CISO’s Guide for Proactive Reporting and DLP with Private AI and Elastic

Jun 5, 2024

The Evolving Landscape of Health Data Protection Laws in the United States

Jun 5, 2024

Comparing Privacy and Safety Concerns Around Llama 2, GPT4, and Gemini

Jun 3, 2024

How to Safely Redact PII from Segment Events using Destination Insert Functions and Private AI API

May 31, 2024

WHO’s AI Ethics and Governance Guidance for Large Multi-Modal Models operating in the Health Sector – Data Protection Considerations

May 29, 2024

How to Protect Confidential Corporate Information in the ChatGPT Era

May 27, 2024

Unlocking the Power of Retrieval Augmented Generation with Added Privacy: A Comprehensive Guide

May 23, 2024

Leveraging ChatGPT and other AI Tools for Legal Services

May 20, 2024

Leveraging ChatGPT and other AI tools for HR

May 17, 2024

Leveraging ChatGPT in the Banking Industry

May 15, 2024

Law 25 and Data Transfers Outside of Quebec

May 13, 2024

The Colorado and Connecticut Data Privacy Acts

May 10, 2024

Unlocking Compliance with the Japanese Data Privacy Act (APPI) using Private AI

May 8, 2024

Tokenization and Its Benefits for Data Protection

May 6, 2024

Private AI Launches Cloud API to Streamline Data Privacy

May 1, 2024

Processing of Special Categories of Data in Germany

May 1, 2024

End-to-end Privacy Management

Apr 29, 2024

Privacy Breach Reporting Requirements under Law25

Apr 27, 2024

Migrating Your Privacy Workflows from Amazon Comprehend to Private AI

Apr 24, 2024

A Comparison of the Approaches to Generative AI in the US and EU

Apr 16, 2024

Benefits of AI in Healthcare and Data Sources (Part 1)

Apr 10, 2024

Privacy Attacks against Data and AI Models (Part 3)

Apr 10, 2024

Risks of Noncompliance and Challenges around Privacy-Preserving Techniques (Part 2)

Apr 10, 2024

Enhancing Data Lake Security: A Guide to PII Scanning in S3 buckets

Apr 9, 2024

The Costs of a Data Breach in the Healthcare Sector and its Privacy Compliance Implications

Apr 5, 2024

Navigating GDPR Compliance in the Life Cycle of LLM-Based Solutions

Apr 2, 2024

What’s New in Version 3.8

Apr 1, 2024

How to Protect Your Business from Data Leaks: Lessons from Toyota and the Department of Home Affairs

Mar 28, 2024

New York's Acceptable Use of AI Policy: A Focus on Privacy Obligations

Mar 26, 2024

Safeguarding Personal Data in Sentiment Analysis: A Guide to PII Anonymization

Mar 21, 2024

Changes to South Korea’s Personal Information Protection Act to Take Effect on March 15, 2024

Mar 18, 2024

Australia’s Plan to Regulate High-Risk AI

Mar 15, 2024

How Private AI can help comply with the EU AI Act

Mar 14, 2024

Comment la Loi 25 Impacte l'Utilisation de ChatGPT et de l'IA en Général

Mar 8, 2024

Endgültiger Entwurf des Gesetzes über Künstliche Intelligenz – Datenschutzpflichten der KI-Modelle mit Allgemeinem Verwendungszweck

Mar 4, 2024

How Law25 Impacts the Use of ChatGPT and AI in General

Mar 4, 2024

Is Salesforce Law25 Compliant?

Feb 15, 2024

Creating De-Identified Embeddings

Feb 6, 2024

Exciting Updates in 3.7

Feb 6, 2024

EU AI Act Final Draft – Obligations of General-Purpose AI Systems relating to Data Privacy

Feb 3, 2024

FTC Privacy Enforcement Actions Against AI Companies

Feb 1, 2024

The CCPA, CPRA, and California's Evolving Data Protection Landscape

Feb 1, 2024

HIPAA Compliance – Expert Determination Aided by Private AI

Jan 23, 2024

Private AI Software As a Service Agreement

Jan 22, 2024

EU's Review of Canada's Data Protection Adequacy: Implications for Ongoing Privacy Reform

Jan 18, 2024

Acceptable Use Policy

Jan 12, 2024

ISO/IEC 42001: A New Standard for Ethical and Responsible AI Management

Jan 12, 2024

Reviewing OpenAI's 31st Jan 2024 Privacy and Business Terms Updates

Jan 12, 2024

Comparing OpenAI vs. Azure OpenAI Services

Jan 9, 2024

Quebec’s Draft Regulation Respecting the Anonymization of Personal Information

Dec 21, 2023

Version 3.6 Release: Enhanced Streaming, Auto Model Selection, and More in Our Data Privacy Platform

Dec 21, 2023

Brazil's LGPD: Anonymization, Pseudonymization, and Access Requests

Dec 13, 2023

LGPD do Brasil: Anonimização, Pseudonimização e Solicitações de Acesso à Informação

Dec 13, 2023

Canada’s Principles for Responsible, Trustworthy and Privacy-Protective Generative AI Technologies and How to Comply Using Private AI

Dec 12, 2023

Private AI Named One of The Most Innovative RegTech Companies by RegTech100

Dec 6, 2023

Data Integrity, Data Security, and the New NIST Cybersecurity Framework

Dec 5, 2023

Safeguarding Privacy with Commercial LLMs

Nov 30, 2023

Cybersecurity in the Public Sector: Protecting Vital Services

Nov 21, 2023

Privacy Impact Assessment (PIA) Requirements under Law25

Nov 16, 2023