Чтобы адаптировать резюме под вакансию или составить сопроводительное письмо, загрузи резюме
описание
Samsung Food connects food, health, and home across the Samsung ecosystem, turning device data into personalized guidance. Its app offers connected, proactive, and personalized experiences, including Vision AI for recognizing food and meals and integrations with Samsung Health and SmartThings.
задачи
Build a failure taxonomy and golden set by hand-grading traces, including sparse-data synthetic profiles, documenting failure categories and their frequency, demonstrating saturation, and freezing holdout traces
Create a binary pass/fail rubric for each taxonomy category, with definitions and examples, and write annotation guidelines
Double-code traces with a second grader, log and resolve disagreements, and maintain a guideline revision log
Create judge prompts for each criterion and validate the judge against human grades on development and holdout sets
Run weekly readouts on helpfulness, relevance, and tone, including a run by the internal owner with the consultant shadowing
Document a revalidation routine triggered by every model or prompt change and run it quarterly
Write a playbook covering grading, taxonomy updates, rubric revisions, judge revalidation, and readouts
Handover the process and test that the internal owner can independently rerun judge validation and a weekly readout
требования
Have run the end-to-end evaluation loop at least once on conversational or generated-text output: trace error analysis, a self-developed failure taxonomy, binary criteria and annotation guidelines, and an LLM judge validated against personal labels
Be able to describe a personally discovered failure mode that had not previously been identified
Be comfortable acting as the arbiter and documenting decisions and reasons for overruling
Understand that the rubric is discovered through grading rather than written upfront
Be able to explain why a judge that agrees 94% of the time may still be useless
Write annotation guidelines that require no interpretation
Be comfortable working in notebooks and spreadsheets; production coding is not required
Nutrition, weight management, and behavior-change expertise are not required; the company has this expertise in-house
Будет плюсом: conversation design, AI product quality, model behavior or policy, content design, human data operations, UX research with strong qualitative coding, RLHF, applied linguistics, product management
условия
The engagement is an approximately 13-week project