Asked at Nvidia

What’s the most relevant GenAI/LLM deployment you’ve recently worked on:

3.What’s the most relevant GenAI/LLM deployment you’ve recently worked on: - What model? Serving stack? - What was the deployment environment: GPU type and count, cloud vs on‑prem, and whether you used a managed service or ran on Kubernetes/bare metal? - What were the latency and throughput goals? 4. How do you debug poor inference performance and what changed or improved after your fix? 5. What part of the GenAI deployment stack are you strongest in, and where would you need ramp-up? 6. What’s your experience with NVIDIA inference tools (e.g., NIM, Triton, TensorRT, TensorRT-LLM, Dynamo)? If you have used any, have you submitted GitHub issues or feature requests (please share links)? 7. What’s your experience with other NVIDIA software or libraries (e.g., CUDA, RAPIDS, NeMo Framework and Microservices, Nsight systems and compute)? If you’ve used any, have you submitted GitHub issues or feature requests (please include links)? 8. Have you contributed to OSS projects? If so, please share GitHub links 9. How would you scope a customer request or technical project? - What questions to ask? - How to focus on top priorities? - How to manage expectations?

AI EngineeringRecruiter screeningSolution Architect, Gen AI
0 answers