Batch inference
Processing many requests together without a latency requirement, typically at ~50% discount on provider batch APIs. The right home for any AI workload a user isn't actively waiting on.
Processing many requests together without a latency requirement, typically at ~50% discount on provider batch APIs. The right home for any AI workload a user isn't actively waiting on.