Assignment 3 — Artificial Intelligence Application
| Item | Due |
|---|---|
| Issue date | Monday, 24 August 2026 |
| Final submission | Saturday, 26 September 2026 at 7:59 am |
Overview
Generative AI has reshaped how software gets built. At the center of this shift are large language models (opens in a new tab) (LLMs), which understand and generate human language with real fluency and nuance. The field has moved past simple prompt-and-response interactions into agentic AI (opens in a new tab): systems that can reason, plan, use tools, and take multi-step actions on their own. Apps can now orchestrate complex workflows, retrieve information and ground responses from external sources, and work across modalities (opens in a new tab).
Coding tools like Cursor, Claude Code, and Codex let developers ship faster than ever, which raises the bar for what counts as a meaningful product. A simple AI wrapper is no longer sufficient; you are expected to build something with real engineering depth and product sophistication.
Modern AI apps offer a few concrete advantages:
- Natural language understanding and generation: LLMs enable apps to not only understand user inputs with remarkable accuracy but also respond in a manner that mirrors human conversation. This grants users a level of interaction that transcends traditional interfaces, fostering more engaging and meaningful experiences.
- Agentic capabilities: AI systems can now autonomously plan, reason, and execute multi-step tasks, from research and data gathering to code generation and workflow automation. This opens up entirely new categories of products.
- Retrieval and grounding: Through techniques like retrieval-augmented generation (opens in a new tab) (RAG), apps can ground their outputs in real, up-to-date data, dramatically improving accuracy and usefulness.
- Multimodal interaction: Modern AI can process and generate text, images, audio, and code, enabling richer, more capable apps.
- Personalization at scale: By leveraging user data and context, apps can deliver deeply personalized experiences that adapt to individual preferences, needs, and behaviors.
This assignment can be done in groups of 3–4 students. If you weren't able to find a group, you'll be randomly assigned one.
This assignment is open-ended by design. We provide milestones so we can grade everyone's app consistently, and to introduce you to the various elements of AI software development in a structured way. We'll also provide related tips, references, and a bit of help to get you started.
The milestones constitute 70% of the assignment's grade. You're free to develop any idea you like. If some milestones don't apply to the app you intend to build, you can petition to replace them with other deliverables and explain why we should agree.
If you intend to do so, submit your petition via email at least one week before the assignment is due. Your petition is subject to approval.
The features you build should fit the aim of your application, ideally seamlessly. For example, LLMs can easily bolt a chatbot onto any app, but a chatbot may not add value to every product; implementing features just because you can may work against you, since "non-value-adding" features distract users from the point of your app.
The expectation for Assignment 3 is much higher than in previous years. With AI-assisted coding tools at your disposal, there's little excuse for a low-quality submission. You should use AI to code, and the bar is substantially higher now. You'll be assessed not just on the product you build, but also on how effectively you leverage AI tools to improve your productivity and output, e.g., what tools you use and why.
To score the remaining 30%, use your creativity to build something that stands out. We won't limit your potential by restricting what you can build; surprise us.
On the engineering side, because AI tools now handle much of the boilerplate, you're expected to push beyond the basics. Practices like maintaining an AGENTS.md (opens in a new tab), writing specs, enforcing formatting and linting, and writing tests are strongly encouraged. They help both you and your AI tools build faster and better. Be prepared to explain how you use AI effectively to build fast while staying on track and maintaining stability.
Objective
In this assignment, your task is to design and implement a web application that uses the capabilities of LLMs and modern AI patterns, persists user data in the cloud, and uses the user's identity in a meaningful way. Aim for significant engineering complexity and system design beyond CRUD. Agentic coding tools, for instance, are highly optimized for code retrieval, coordinating multiple async jobs, and managing memory: features that require solid engineering that pure vibe coding can't achieve. Go beyond simple prompt-and-response interactions and incorporate modern AI patterns (see Overview).
Approach this assignment with the mindset of an entrepreneur. You own every decision, and each one you make, from design to engineering, directly shapes the "success" of your product.
Use this assignment to showcase your product sense and engineering capabilities. Consolidate what you learned from Assignment 1 (on Product Design) (opens in a new tab) and Assignment 2 (on identifying innovations, and gaps in the market (opens in a new tab)) to build a fuller product.
Remember, your goal isn't to do a lot of work; it's to use this opportunity to make a difference.
Please read the entire assignment before deciding what to develop. It should give you a clearer picture of how the parts fit together, and a better understanding of the strengths and weaknesses of an AI application. The assessment scheme is at the end of this handout.
Phase 0: Introduction
Background
A new wave of applications has emerged over the past few years, built around generative AI and LLMs; built on top of providers like OpenAI (ChatGPT, DALL-E) and Stability AI (Midjourney), who demonstrated what AI-generated content and interaction could look like.
Big tech has fully joined in: Google, Microsoft, and Meta have all entered the AI race with their own models and chatbots, dedicating significant resources to AI research as they integrate it across their existing products.
The shift isn't limited to established players. A growing ecosystem of AI startups, particularly in copywriting and customer support, is reshaping how businesses communicate with users. 35% of Y Combinator's Summer 2023 batch were AI startups (opens in a new tab), and that share has only grown since.
It's an exciting time to be a software engineer.
New AI Landscape
Within the generative AI startup landscape, there are a few common categories of products:
- Text: Create and manipulate textual content. Tools that can draft articles, generate product descriptions, and even assist in creative writing. Examples: Copy.ai (opens in a new tab), Jasper (opens in a new tab), Hypotenuse AI (opens in a new tab).
- Image: Generate, modify, or enhance images. Tools for artists, designers, and photographers to generate artwork, edit photos, and visualize ideas. Examples: Midjourney (opens in a new tab), Runway ML (opens in a new tab), Pebblely (opens in a new tab), Adobe Photoshop.
- Audio: Compose music, generate sound effects, and even mimic specific voices.
- Code: Assist in software development tasks, including generating code snippets, offering coding suggestions, and even automating parts of the coding process. Examples: GitHub Copilot (opens in a new tab), Sourcegraph Cody (opens in a new tab).
- Chatbot: Create conversational agents powered by generative AI. These chatbots can engage in natural conversations, answer queries, and provide support based on custom data. Examples: Mendable (opens in a new tab), Chatbase (opens in a new tab), Glean (opens in a new tab).
- Design: Design your brand, logo, websites, presentations, and marketing collateral with a prompt. Examples: Framer AI (opens in a new tab), Designs.ai (opens in a new tab), Uizard (opens in a new tab), Gamma (opens in a new tab).
- Video: Manipulate and create video content. These tools can be used for video editing, special effects, and even automated video creation. Examples: Synthesia (opens in a new tab), Lumen5 (opens in a new tab).
- Data and Analytics: Analyze and generate datasets for testing and simulation purposes. Query data using natural language. Examples: Defog (opens in a new tab).
- Agents: Create virtual agents powered by generative AI. These agents can emulate human interactions and assist with tasks like scheduling, information retrieval, and more. Examples: Coze (opens in a new tab), Flowise (opens in a new tab).
- Gaming: Create dynamic game environments, generate levels, adapt game mechanics based on player behavior, NPCs can engage in personalized conversation with players.
While most companies are part of the AI "gold rush", some prefer to follow the saying "during a gold rush, sell shovels". These "shovel" companies build services around LLMs, selling API access to LLMs and platforms to make it easier to build AI products:
- APIs: Access to the LLMs hosted on the cloud. Examples: ChatGPT & GPT-4o by OpenAI, Claude by Anthropic, Command by Cohere.
- Toolchains: Simplify common LLM-related operations. Examples: LangSmith by LangChain, Cohere platform, Humanloop.
- Vector databases: Stores data in a format that enables semantic information retrieval and long-term memory for LLMs. Examples: Supabase Vector, Pinecone, Weaviate.
Other lists to find AI products:
- Product Hunt (opens in a new tab) (everything is mostly AI-powered nowadays)
- 152 Fun AI Tools You've Never Heard Of (opens in a new tab)
- There's An AI For That (opens in a new tab)
Phase 1: Product Strategy
Idea Generation, Problem Space
As an aspiring entrepreneur in CS3216, you probably have a mind full of big ideas. But before getting too attached to any one of them, ask yourself: "Does my application solve a real problem for its users?" A cool concept or the latest technology doesn't guarantee usefulness: if a plainer but more practical solution makes users' lives easier, they'll stick with that instead. Solve a problem people care about, solve it well, and users will become advocates who spread the word for you.
When deciding on a problem space, it's also worth considering market attractiveness: market size and expected growth rate can help you gauge the potential impact of your product.
Guides on estimating market size are all over the internet; a useful one is pear.vc/market-sizing-guide (opens in a new tab).
Describe the problem that your application solves.
Competitive Landscape
Understanding the competitive landscape gives you critical insight into existing players, their offerings, and their strengths and weaknesses; helping you to anticipate challenges and find a way to deliver unique value.
The AI product ecosystem is crowded, and new AI companies emerge constantly. That's unlikely to change any time soon.
AI has also given rise to "thin wrapper" products that rely heavily on existing frameworks and APIs while adding minimal value. Do NOT build one of these as they are easy to clone and don't hold up long-term.
To understand the competitive landscape, a competitive analysis can help. Useful questions include:
- Is the market you're entering competitive, and will you be able to capture meaningful share?
- How much market share does each competitor have?
- What are your competitors' competitive advantages?
Good to know: You may consolidate your findings into a Competitive Profile Weighted Matrix to compare your product against competitors (not graded). This is commonly used in strategic management: identify the key success factors that determine success in your market, and use them to compare your competition.
List down your 3 closest competitors and their pros and cons. Explain how your product is better.
Product Capabilities
You should choose an application that fully uses the potential of the technology you've chosen; execution matters. It's on you to justify why LLMs are the right fit for your application's objectives.
Your application must also have user authentication, used meaningfully. We're looking for products that are more than one-page tools. From a product standpoint, authentication should let you build personalization features and force you to think about protecting your users' data. Explain why authentication alongside LLMs helps you meet your application's objectives.
A common pitfall for engineering-focused students is bias toward technologically complex products, without considering whether the product actually serves the problem. Ask yourself: does your product have the capabilities to succeed in your chosen problem space?
All teams must first post their proposed product on Coursemology under the Assignment 3 topic for the teaching team's approval before starting. The teaching team will evaluate ideas carefully to make sure they push beyond what can be trivially built with AI tools alone.
Describe your application briefly. List its objectives and the associated (major) user stories.
Moat
In product strategy, a "moat" is a sustainable competitive advantage that protects a product's market share and profitability: like a moat around a castle, it makes it harder for competitors to breach.
The moat question matters more now than in previous years. With AI-assisted coding tools, cloning an app is easy. If a competitor can replicate your product in a weekend using Cursor or similar tools, you don't have a moat. Think hard about what makes your product defensible.
A moat can take various forms, including (but not limited to):
- Brand
- Technological Innovation
- Economies of Scale
- Network Effects and Data Advantages
- Proprietary Datasets or Domain Expertise
- Investor Confidence
- Ethics and Responsible AI
- Customisation and Personalisation
A few examples: Google's moat is scale: Google Search still dominates the search market.

One of Apple's moats is its tightly integrated ecosystem: iPhone, Mac, AirPods and services like iCloud are all tightly integrated together into a single ecosystem, which makes switching away difficult for consumers. Disney's moat is its intellectual property: owning characters from Mickey Mouse to Star Wars and Marvel is a competitive advantage that's hard to erode.
Closer to home, ChatGPT has a few moats of its own:
- First-mover advantage: strong brand recognition as the pioneering product, which is hard for later entrants to match. Everyone still says: "let me ask GPT"
- Developer-friendly APIs: easy integration into other applications and services without deep AI expertise.
- Demand scale: widespread usage feeds data collection and training improvements, enhancing the model's performance over time.
For AI startups especially, establishing a moat is essential to stand out in a fast-moving market.
What's your secret sauce / moat? Elaborate on your strategy to prevent competitors and big players from cloning your app and its features. Given how easy it is to clone apps with AI tools, what makes your product defensible?
Phase 2: Go-To-Market
Product Lifecycle & Product-Market Fit
The product lifecycle describes the stages a product goes through, from launch to eventual sunset. Your product sits in the introduction stage, a period that demands actively engaging prospective users to become your first users.
The introduction stage isn't primarily about achieving optimal product-market fit, but it lays the foundation for it: this is when you collect feedback, understand user preferences, and make early adjustments.
A user acquisition plan should help you reach more users in your target segment, and it might shape which features you build. For example, you could design features that reward users for referring friends, giving them a reason to spread the word even as they get value themselves.
Describe the strategies, channels, and tactics you'll use to reach and convert potential users, and how your understanding of your target users shapes that approach.
Approach this milestone by describing your strategy or plan for bringing this product to market, using whatever resources you have.
Describe your target users. Explain how you plan to acquire your target users.
Scoping
With your problem and target users clear, it's time to decide what features actually go into the product you submit for Assignment 3.
Time-to-submission is short. Like real companies navigating deadlines and customer expectations, you'll need to optimize your time to deliver a simple, lovable, and complete Minimum Viable Product (MVP).
This is where scoping, carefully selecting which features to include, matters. Justification matters too: a good scoping decision requires understanding user needs, potential impact, and overall project goals. Useful questions include:
- User Experience: How does each proposed feature address a specific user pain point? What value does it add to the experience?
- Resource: What are the time, effort, and cost implications of including each feature? Are there technical limitations or dependencies?
- User Impact: How many users benefit from each feature? Is it essential for a large share of your user base, or a niche requirement?
- Competitive Advantage: Does including this feature give you a competitive edge? How does it compare to what similar products offer?
Since you're building something new, your initial scope matters a lot: ship a first set of features as an MVP, release it, gather feedback, and iterate, the same logic behind agile development, which lets teams respond quickly to changing needs.
As you prioritize features, balance planning and estimation: breaking a project into actionable tasks, foreseeing challenges, allocating resources, and estimating timeframes are all skills worth building here.
You can divide your 3–4 weeks into sprints (opens in a new tab), using each one to better estimate the next. Since this is a short project, don't let strict adherence to sprints get in the way of shipping: use your judgment.
List down the features that should go into the MVP (your assignment deliverable). How did you decide on them? What are future features and expansions you can think of?
Business Model
Whether your product succeeds often comes down to pricing, among other things. What counts as "success" shifts as a product moves through different stages of growth: sometimes it's about maximizing profit, other times about maximizing market share or user count.
There are generally three ways to price:
- Based on costs: how much does it cost to produce and maintain your product, and what profit margin are you aiming for?
- Based on value: what benefit does the product provide users, and how much is that worth to them?
- Based on competition: what do competitors charge, and how does your product compare?
Having a monetization strategy early helps you evaluate the right price for your current stage: for instance, an early-stage product might prioritize market share and time to break even over margin.
A couple of resources that may help you think through pricing:
- The Ultimate Guide To SaaS Pricing Models, Strategies & Psychological Hacks (opens in a new tab) covers seven common ways SaaS products are priced today.
- Mapping the Generative AI landscape (opens in a new tab)'s section on "Looking into the future — Gen-AI revenue models."
Warning: AI products need different pricing models from traditional SaaS. Many AI companies operate at negative margins today, betting that costs will fall or that they can raise prices later. Usage-based pricing (per API call, per token, per task) is increasingly common. Think carefully about how AI inference costs affect your pricing strategy, and whether your business model is sustainable: weigh flat subscriptions, usage-based pricing, and freemium against your variable compute costs.
"Success" doesn't have to mean dollars. If not revenue, what metric matters for your product? For a social chatbot, for instance, message volume might be a better signal.
Come up with a monetization and pricing strategy (e.g., tiers and features). Explain why you think this pricing strategy is suitable for your target users and problem space. Explain the factors that influenced your pricing decisions, such as production costs, perceived value, competition, and AI inference costs. It would be useful here to consider possible revenue streams of your product and how your cost structure scales with usage.
Phase 3: Artificial Intelligence Integration
Introduction to Large Language Models
LLMs are a type of AI designed to understand and generate human-like text based on the input (prompt) they receive. They're built using deep learning, particularly a neural network architecture called transformers, the foundation behind well-known models like OpenAI's GPT (Generative Pre-trained Transformer).
LLMs have billions of tunable parameters learned from massive amounts of text, allowing them to perform a wide range of language tasks:
- Text generation: generating coherent, contextually relevant text from a prompt.
- Language translation: translating between languages.
- Summarization: producing concise summaries of longer text.
- Question answering: extracting relevant information to answer questions.
- Text classification: classifying text into predefined categories.
- Sentiment analysis: determining emotional tone.
- Chatbots and conversational agents: simulating human-like conversation.
- Code generation: generating code from high-level descriptions.
- And so much more…
The training process involves exposing the model to vast amounts of text and having it predict the next token in a sequence; this is how it learns grammar, syntax, semantics, and some world knowledge.
Compared to traditional machine learning, LLMs:
- Handle more diverse, complex tasks through a single, flexible natural language prompt, instead of being limited to narrow tasks like classification or regression.
- Are readily accessible via APIs, without requiring deep ML expertise, and require no new syntax to use.
- Are trained at massive scale on unlabeled data, rather than small labeled datasets.
- Support few-shot and zero-shot learning: performing new tasks with minimal or no examples, which matters when labeled training data is expensive or hard to get.
- Understand nuance: idioms, metaphors, and cultural references that are hard for traditional models to grasp.
Resources:
- A Compact Guide to Large Language Models (opens in a new tab) by Databricks.
- How ChatGPT works (opens in a new tab) by Sau Sheong Chang, a good look under the hood at the concepts behind ChatGPT. He's kindly given CS3216 students free access to this article.
Explain how you are using LLMs in your product and why LLMs are a good approach to meet the product's objectives.
Prompt Engineering/Design
A prompt is the input or instruction you give a language model to guide its behavior and get the output you want.
Prompts can be as short as a sentence or as long as a paragraph. Since prompts are the main interface between users and LLMs, writing a good one matters: word choice, phrasing, and clarity all shape the quality of the output.
Elements of a Prompt
A prompt typically includes several elements that help guide the response or output generated by the LLM. Here are some common elements of a prompt:
- Topic/subject: The topic or theme of the prompt provides a broad area of focus for the response. For example, a prompt might ask a language model to generate text about a specific topic like "climate change" or "artificial intelligence."
- Task/goal/purpose: The task or goal of the prompt specifies what kind of response is desired. Is it to persuade, inform, entertain or something else? For example, a prompt might ask a language model to generate a persuasive essay on a particular topic or to write a short story that meets certain criteria.
- Target audience / perspective: The target audience of the prompt identifies the intended audience for the response. For example, a prompt might ask a language model to generate text that is suitable for a general audience or for a specific age group. For example, "Respond as if you are a teacher providing advice to a student."
- Tone/style: The tone or style of the prompt can influence the tone or style of the response. For example, a prompt might ask a language model to generate text that is formal, informal, humorous, or serious.
- Context / background information: Providing relevant background information helps set the stage for an informed response.
- Length/format: The length or word count of the prompt can specify the amount of content that is expected in the response. For example, a prompt might ask a language model to generate a response that is between 500 and 1000 words.
- Specific instructions/guidelines: The prompt may include specific instructions or guidelines that must be followed in the response. For example, a prompt might ask a language model to generate text that includes specific keywords or phrases, or that adheres to a particular format or structure. For example, "List 3 reasons why..."
Prompt engineering is a broad topic and cannot be covered sufficiently within this writeup. Here are some recommended resources for learning more about prompt engineering:
- Best practices for prompt engineering with OpenAI API (opens in a new tab) by OpenAI. One-pager that goes straight to the point, with examples.
- GPT best practices (opens in a new tab) by OpenAI. A longer list of techniques and best practices.
- ChatGPT Prompt Engineering for Developers (opens in a new tab) by DeepLearning.AI. This video-based course takes around an hour to complete and Yangshun has written a summary of the important points: Part 1 (opens in a new tab) & Part 2 (opens in a new tab).
- Prompt Engineering Guide (opens in a new tab) by DAIR.AI. Well-rounded useful guide that covers Prompt Engineering and various topics related to LLMs.
- Prompt Engineering Guide (opens in a new tab) by Brex. Tips and tricks for working with LLMs like OpenAI's GPT-4.
- Prompt Engineering (opens in a new tab) by Cohere.
Tokenization
Tokens are the units, subwords, words, or characters that a model processes and generates. Both input and output tokens count toward a single API call, which affects cost (most providers charge per token), latency, and whether you stay within the model's context window (e.g., 128k tokens for gpt-4o-mini, 1m+ for gemini-1.5-flash).
Further reading:
- How to count tokens with tiktoken | OpenAI Cookbook (opens in a new tab)
- Tokens and Tokenizers | Cohere Docs (opens in a new tab)
Fine-tuning
Fine-tuning lets you train a model on examples outside the prompt itself, rather than relying only on "few-shot prompting". Fine-tuned models tend to offer better results, can learn from more examples than would fit in a prompt, and produce shorter, cheaper, lower-latency requests since fewer examples need to be in the prompt.
In general, fine-tuning involves: preparing and uploading training data, training the fine-tuned model, then using it. It's specific to particular models: fine-tuning is supported by various providers for models such as Llama, Mistral, Gemma, and Qwen.
Further reading:
Give two to three examples of prompts you used and explain how you designed them to be effective. What techniques did you use to improve the effectiveness of your prompts?
Using the Right Model for the Job
Not all models are equal: each has its own strengths, limitations, and tunable parameters. The key is finding the model that best fits your application. Compare quality, reliability, latency, context window, multimodal and tool support, cost, and provider constraints. Test at least two alternatives on representative examples and explain the trade-offs instead of assuming the largest model is best.
Popular LLMs and Providers
- OpenAI (opens in a new tab): a pioneer in LLMs, offering state of the art models in their GPT lineup.
- Claude (opens in a new tab) by Anthropic: Anthropic's model family Fable/Opus/Sonnet/Haiku, known for reasoning, coding, and safety.
- Llama (opens in a new tab) by Meta AI (opens in a new tab): open-weight models, self-hostable or available via Groq, Together AI, AWS Bedrock.
- Hugging Face (opens in a new tab): a community platform hosting thousands of open models, datasets, and deployment tools.
- Gemma & Gemini by Google AI (opens in a new tab): Google's multimodal family, with context windows up to 2 million tokens, available via Google AI Studio or Vertex AI.
- Nemotron by NVIDIA (opens in a new tab) : Providing optimized inference microservices (NIMs) and frameworks like NeMo for deploying high-performance AI models.
- Grok (opens in a new tab) by xAI (opens in a new tab) : The Grok family of models, known for real-time access to platform data and advanced reasoning capabilities.
- Command (opens in a new tab) by Cohere (opens in a new tab): the Command family (Command R, R+), optimized for RAG and tool-use.
- Qwen (opens in a new tab) by Alibaba Cloud (opens in a new tab) : The Qwen series of models, highly capable in multilingual tasks, coding, and mathematics.
- GLM (opens in a new tab) by Zhipu AI (opens in a new tab) : The General Language Model (GLM) series, known for strong performance in bilingual (Chinese/English) and multimodal tasks.
- Kimi (opens in a new tab) by Moonshot AI (opens in a new tab): The Kimi series of models, known for long-context windows and strong reasoning capabilities.
Vercel AI Playground lets you compare outputs from different providers using the same prompt.
When selecting an LLM and provider, prioritize application requirements, training data, context window capacity, and cost per token. Note that smaller, domain-specific models often deliver superior cost-to-performance ratios compared to general-purpose alternatives.
LLM APIs generally aren't free, though costs should be low for an assignment's usage volume. If you're having trouble getting API access, please email the teaching staff.
Model Settings
Beyond the prompt itself, most models expose settings for finer control. Know that not all models/providers support these parameters.
- Max tokens: the maximum number of tokens to generate.
- Temperature: controls creativity and diversity. Higher temperature means more creative, more varied (and less predictable) output; lower means more consistent, accurate output. Use low temperature for factual tasks and higher for creative ones. Reasoning models don't expose this parameter.
- Top-P: limits the probability mass considered when sampling, e.g., Top-P of 0.75 only considers tokens covering the top 75% of probability. Customize either temperature or Top-P, not both.
- Frequency/presence penalties: penalize tokens based on how often they've already appeared, to reduce repetition.
Justify your choice of LLM and provider by comparing it against at least two alternatives. Explain why the one you have chosen best fulfills your needs. Elaborate on your choice of model parameters.
Evaluating effectiveness
Once your LLM integration is up and running, set up a lightweight evaluation pipeline to check whether your prompting or fine-tuning is actually working across test examples. Broadly, there are two families of Natural Language Generation metrics:
- N-gram-based metrics measure word or token overlap against the reference. E.g., BLEU, METEOR.
- Model-based metrics use a neural model to measure similarity against the reference. E.g., BLEURT, BERTScore.
N-gram metrics suit tasks where precise wording matters; model-based metrics suit more open-ended generation. There's no single correct metric: use whatever's standard for your task (BLEU for translation, ROUGE for summarization), and note that a single test input can have multiple valid reference answers.
Further Reading:
- Neural Language Generation (opens in a new tab), Chris Manning, Stanford: Slides 51 to 65
- A Comprehensive Assessment of Dialog Evaluation Metrics (opens in a new tab)
- Holistic Evaluation of Language Models (opens in a new tab), Percy Liang et al. (2023), Stanford NLP Group
- Instruction Tuning for Large Language Models: A Survey (opens in a new tab)
Modern AI Engineering Patterns
Building a meaningful AI product today goes beyond sending a prompt and displaying the response. A few patterns are worth exploring where they add value to your product (see Overview for agentic AI):
- RAG (opens in a new tab): supplementing the model with relevant information retrieved from your own base before generating a response, grounding output in real data and reducing hallucinations.
- Tool calling / function calling (opens in a new tab): letting the model invoke external tools or APIs (search, databases, calculators, code interpreters) as part of its reasoning.
- Structured outputs (opens in a new tab): constraining output to a specific format (JSON, typed schemas) for reliable downstream processing.
- Multimodal input/output: processing and generating across text, images, audio, video, and code, increasingly easy to incorporate, and worth considering over text-only output.
It's worth distinguishing chat interfaces (stateless back-and-forth), agents (autonomous systems that plan, reason, and act using tools, see Overview), and workflows (structured, multi-step processes orchestrating multiple AI operations). Consider how your product can go beyond a simple chat interface.
Building Agentic Systems
As introduced in the Overview, agentic systems represent the next evolution of AI applications, going well beyond simple prompt-response interactions. Key components include:
- Tools: functions or APIs the agent can invoke to interact with the world (web search, database queries, file operations, API calls).
- Memory: mechanisms for maintaining context: short-term (conversation history) and long-term (persistent knowledge).
- Planning/reasoning loop: the core loop where the agent decides what to do next based on goals, observations, and available tools.
- Orchestration: coordinating multiple agents or sub-tasks toward a complex goal.
If you're building an agentic system, research and compare a few agent SDKs and frameworks, and explain why you chose the one you did.
Resources:
- 12-Factor Agents — Principles for building reliable LLM applications (opens in a new tab): A set of principles for building reliable, production-grade agentic systems. Highly recommended reading.
Describe the AI interaction patterns you are using in your product (e.g., RAG, tool calling, structured outputs, agent loops, multi-step workflows). Explain why these patterns were chosen and how they contribute to your product's objectives. If you evaluated multiple frameworks or SDKs, explain your selection process.
LLMOps and Evaluation
Just as DevOps transformed how software is deployed and maintained, LLMOps (Large Language Model Operations) is an emerging discipline for managing AI systems in production. A key part of it is evaluation: systematically measuring whether your AI system performs as expected. You should:
- Create an evaluation dataset: curate representative inputs and expected outputs (or quality criteria) for your key use cases.
- Implement evaluation strategies: build eval pipelines that automatically assess output quality, automated scoring, LLM-as-judge, or human evaluation.
- Iterate based on results: use eval results to guide prompt engineering, model selection, and system design.
Resources:
- Demystifying Evals for AI Agents (opens in a new tab) by Anthropic, a good guide to evaluating agentic AI systems.
Describe the evaluation dataset and strategy you created for your AI system. How do you measure whether your system is performing well? Show examples of your eval results and explain how they influenced your development decisions.
Loop Engineering
This is an emerging term. It entails designing the goal, shared context (like your AGENTS.md), hooks, and checks around a coding agent so it can act, observe the results of its own work, and adjust across many steps with minimal re-prompting. It's the practice underlying the coding tools like Cursor, Claude Code, and Codex.
These components: AGENTS.md, specs, and tests, they give the agent something concrete to check itself against instead of needing you to verify every step. That said, a well-designed loop doesn't remove the need for human review: teams that accept AI-generated code unquestioningly can quietly build up "comprehension debt," not fully understanding code they remain responsible for. See IBM's overview of loop engineering (opens in a new tab) for a breakdown of loop components (hooks, subagents, persistent state), and Anthropic's Claude Code best practices guide (opens in a new tab) for concrete patterns, like giving the agent a test or build step it can run to check its own work.
Production Optimization
Running AI systems in production comes with its own performance and cost challenges. Techniques worth applying include:
- Prompt caching: reusing cached responses or partial computations for repeated or similar queries.
- Batch calls: grouping requests together to improve throughput.
- Minimizing latency: streaming responses, using smaller/faster models for simpler tasks, parallelizing independent operations.
- Cost management: monitoring token usage, right-sizing models per task, and building fallback strategies.
Describe the production optimization techniques you implemented. What was the impact on latency, cost, or user experience? Provide concrete metrics where possible.
Safety and Security
AI applications carry their own safety and security considerations; this is a graded component, not optional. Key areas:
- Prompt injection: guarding against adversarial inputs trying to manipulate your system into unintended outputs or actions, especially critical for agentic systems that can take real-world actions.
- Output validation: ensuring AI-generated output is safe and appropriate before it reaches users or triggers actions.
- Data privacy: protecting user data sent to or processed by AI models, including retention and third-party API usage.
- Rate limiting and abuse prevention: preventing excessive or malicious use of your AI system.
Identify the key risks and safeguards in your AI system. Describe at least two specific safety or security measures you implemented (e.g., prompt injection guards, output filtering, rate limiting). Explain the threat model and how your measures address it.
Other Resources
Courses and Tutorials
- DeepLearning.AI (opens in a new tab) by Andrew Ng provides a series of short, practical courses on generative AI (opens in a new tab), video-based with hands-on Jupyter Notebook exercises.
- Creating a private QA over local documents application using Llama-2 (opens in a new tab) by Sau Sheong Chang.
Tools
LLMs come with some friction: unstructured output, preserving context across completions, boilerplate code. A few tools that help:
- LangChain (opens in a new tab): abstractions for common LLM operations, loading/transforming/storing data, chaining operations, remembering interactions, acting on output.
- Vercel AI SDK (opens in a new tab): build AI-powered UIs in JavaScript and TypeScript; makes streaming chat interfaces easy.
- TypeChat (opens in a new tab): get well-typed, structured responses from LLMs without parsing output yourself.
- Web LLM (opens in a new tab): run LLMs and LLM-based chatbots directly in the browser, for free (device requirements apply).
- LlamaIndex (opens in a new tab): connect and query your own data sources for LLM apps.
For production use (possibly overkill for this assignment, but worth knowing about):
- LangSmith (opens in a new tab) (visualize, debug, and improve LLM apps)
- Humanloop (opens in a new tab) (monitoring, A/B testing, a collaborative prompt workspace, and fine-tuning).
Open Source LLMs
Platforms like Hugging Face (opens in a new tab) host open-source LLMs built by the community to mirror proprietary models. They're easily fine-tunable and lightweight, not as powerful as commercial LLMs, but a free, capable option for many use cases. Note that while these open models are generally fine-tunable, non-consumer-hostable open weights models like GLM or Kimi models are ideally accessed via API. Another free option worth knowing about: NVIDIA NIM (opens in a new tab).
A few popular models and guides to fine-tune them:
Fine-tuning these models works best with access to a high-performance GPU (e.g., NVIDIA RTX 3090, RTX 4090, H100 Tensor Core), or a cloud platform like Lambda Labs (opens in a new tab), Paperspace (opens in a new tab), or Google Colab Pro/Pro+ (opens in a new tab).
You can also apply for free cloud credits through Google Cloud's Research and Education program (opens in a new tab) (application-based, no guarantees, but worth a shot).
Phase 4: Design
Design here covers branding, technology/engineering, and the product.
Branding your Product
Every product needs a name and, usually, a logo; together with color schemes and fonts, these form the persona people associate with your product. That persona is your brand, and it's rarely accidental; it's carefully crafted and tested.
Think of McDonald's golden arches: instantly recognizable, and long since a stand-in for consistency, quick service, and affordability. A strong brand can even become part of the language: Singaporeans still say "maggi mee" for instant noodles, and much of the world says "google it" instead of "search the web." That's the power of a good brand: it builds trust and revenue at once.
Choosing a name is often harder than it looks. Consider:
- Length of the name. Will it be too hard to remember?
- Reproducibility of the name by spelling. If A were to hear B say "I use Whatchamacallit to build this website," would A be able to type it and find it on search engines?
- Domain name availability. Some may even get clever with domain hacks (opens in a new tab), e.g., del.icio.us (opens in a new tab), but can it be easily shared verbally?
- Social handles availability. Is @mynextstartup available on major social media sites, e.g., Instagram, X, Threads, TikTok, etc.? If you're not quick to claim the handle, adversaries may claim it first, and you might have to either buy it from them, or file a dispute.
- Similarities with other brands. Could people think your product Zwitter is just a cheap clone of Twitter?
- Cultural references. Did you know the toothpaste brand Darlie was originally known as Darkie? It was changed in 1989 due to "darky" or "darkie" being considered a racial slur.
- Possibility of being caught in web filters. In April 1998, "shitakemushrooms.com" was blocked by DNS filters because it contained the word "shit". This is known as the Scunthorpe problem (opens in a new tab). If you want to see more examples, check out the article "30 Unintentionally Inappropriate Domain Names (opens in a new tab)".
These are just a fraction of factors that one should think about when choosing names. Different factors apply when you are choosing logos and color schemes. Think about how they'd look if they are printed on different materials. Do the logos have any resemblance to a potentially derogatory subject, seen from any angles? How easy is it to be reproduced by hand? Can it be a vector (so that it is scalable on different media)? How'd it look if it was black-and-white, or against a busy background? Think about Apple's first logo (opens in a new tab).
There are many logo generators out there, but after reading this far, you'd probably realize that a good brand makes for a good logo; it is not standalone, it is an integral part of a brand. To make a good brand identity (name, logo, fonts, color schemes, etc.), you should reflect on the persona that you wish your company or product to convey to your users, then work forwards. Use Pinterest (opens in a new tab), Dribbble (opens in a new tab), Midjourney (opens in a new tab), Coolors (opens in a new tab), Google Fonts (opens in a new tab), or even the many logo generators to brainstorm for inspirations. Then choose a comfortable medium to sketch, e.g., pen and paper, Figma, Microsoft Paint, etc., to stitch everything together. Test your brand identity prototypes with your friends and family, and see if they can guess what app you're building just by seeing your logo.
Notably, a good brand identity requires tons of research and testing. Schools of design talk about branding for one whole semester. Due to the density of this assignment, you are expected to build a decent enough brand identity, i.e., with proper considerations. You shouldn't dwell too long on this milestone. It is maybe even best to first build your MVP with a codename, only after then collaboratively think about the best logo to slap on your app. If you did well enough, you may even earn brownie points for coolness!
Come up with a product name and create an attractive logo. Explain the meaning behind the name, the alternatives you've considered, and why this was chosen.
Technology Stack
A functional product is backed by a technology stack: the collection of tools and frameworks supporting its business logic (opens in a new tab).
Newer technologies tend to get attention because they fix pain points in current "standard" tools (Svelte (opens in a new tab)'s syntax vs. React (opens in a new tab), Deno (opens in a new tab)'s package management vs. Node (opens in a new tab), Vite (opens in a new tab)'s speed vs. webpack (opens in a new tab)), but they also tend to have less community support and less battle-testing. Some fade quickly; others stick around. The Stack Overflow 2025 survey (opens in a new tab) is a useful reality check on hype vs. staying power.
Before picking new tech, weigh:
- Your team's familiarity with the technologies
- How well it scales with your users and (potentially increasing) business logic
- The maturity of the ecosystem; How mature its ecosystem is (can you get help when something breaks?)
- Your reliance on the technology; How reliable its maintainers are, and how easy it'd be to switch away from later.
Generally, it's better to pick tools most of your team already knows and start building, and failing fast. That said, exercise judgment: The Browser Company built the Windows version of Arc Browser (opens in a new tab) by extending Swift's (then quite barebones) Windows support for SwiftUI (opens in a new tab) instead of using the more obvious choice, Electron. Their CTO talks about that decision, a good example of challenging the default for their specific use case.
Your users won't care what's under the hood, but every choice comes with trade-offs: what you save today, you may pay for tomorrow as technical debt (opens in a new tab).
A few domains worth considering as you pick your stack:
- Front-end: what users interface with directly. Examples: React, Vue.js (opens in a new tab), Next.js (opens in a new tab), Svelte, etc.
- Styling: UI component libraries that make it easier to include elements like buttons, menus, and cards, or build from scratch if your design calls for it. Examples: Tailwind CSS/UI, Bootstrap, Bulma, MUI, etc.
- Data persistence: SQL (e.g., MySQL (opens in a new tab), PostgreSQL (opens in a new tab)) for relational data, NoSQL (e.g., MongoDB (opens in a new tab), Redis (opens in a new tab)) otherwise, or higher-level options like Firebase (opens in a new tab) or Supabase (opens in a new tab). Your schema choice directly affects scalability and query performance, and migrations aren't always reversible; choose carefully.
- Authentication: OAuth is usually the easiest path, given how simple it is to integrate; some sensitive apps (banks, government systems) roll their own for more control. Examples: self-rolled, Firebase Authentication, Auth0.
- Back-end server: where your business logic lives, communicating with third-party APIs, databases, and caching servers. Examples: Go, Django, Ruby on Rails, Express, etc.
- Hosting: from simple static hosting (e.g., Cloudflare (opens in a new tab), Netlify (opens in a new tab), GitHub Pages (opens in a new tab)) to batteries-included platforms (e.g., Vercel (opens in a new tab), Firebase (opens in a new tab)) to raw servers (e.g., Google Cloud (opens in a new tab), AWS (opens in a new tab)), a trade-off between complexity, flexibility, and cost.
- CI/CD: platforms that automate code checks and deployment, helping your team spot errors early and deploy consistently. Examples: CircleCI, Travis CI, GitHub Actions.
- Miscellaneous services: whatever your app needs beyond the basics: LLM-backed services, maps, search, computer vision, etc. The less you need to build yourself, the faster you'll ship.
If any of your chosen services aren't free, factor those costs into your business model; your revenue needs to cover operational costs well before you're profitable.
Explain choice of technologies for the following: UI, Database, Web Server, Hosting, Authentication, etc. and the alternatives you've considered.
User Experience
User experience (UX) is very different from the user interface (UI). A pretty UI does not guarantee a good UX at all. Sometimes, a cool-looking UI can be a disaster because of poor UX. Sites like Craigslist still have a ton of daily active users despite their UI! 😵
UX encompasses all aspects of the end-user's interaction with a product. There are many components that comprise a good UX, but minimally the application should allow the users to do what they want with minimal fuss. UX goes far beyond giving users what they say they want or providing a checklist of features. In order to achieve a high-quality user experience in your application, careful thought must be put into its functionality, interaction design and interface design.
UX is about research. It requires understanding of your user. Not just when they are using your product, but also the environments around the time of usage, the different agents surrounding the product and its implications, and ultimately their objectives. If you are building an app that is used in rowdy buses or busy places, can your users tap on those small buttons (this is known as a situational disability (opens in a new tab))? If you are building an app for teachers to input grades, how do you think it will impact the parents of the teachers, and the students (they are known as primary, secondary, and tertiary users (opens in a new tab))? There are no right questions, just as there are no right answers. Talk to your users, and empathize with their objectives. Return to the drawing board and devise a prototype. Test it with them, collect feedback, be surprised, and iterate. Just like software engineering, product development is an iterative (read: agile) process.
You are not your own user. Repeat it. You are not your own user. It doesn't matter if you think you fall into the target audience of your own app, you are still not, and never will be, your own user. It is therefore important to distance oneself when researching user behaviors: be ready for surprises, and be humble to counterintuitions. If you're interested, the field that heavily studies these subjects in the context of computers is called Human-Computer Interaction (HCI) (opens in a new tab).
UX is not just the job of the UI designer. Just like a good UI, you will know if an application's UX is good or bad. Any team member can contribute to your UX design by using it and providing feedback and suggestions. Ask your friends to use it as well to gather more feedback and ideas.
You can probably guess now that UX is somewhat related to branding. Diction that prefer colloquialism over specificity may convey a relaxed, friendly sentiment (In 2018, Gojek revamped their app to relate better to the locals (opens in a new tab), and Don Norman talked about designs that make people happy (opens in a new tab)). Smooth transitions and tidy designs may convey a sense of stability and trust (known as the aesthetic-usability effect (opens in a new tab)). Animations and choice of materials (gradients, acrylic, shadows, elevations, etc.) give a sense of space and control, that the user knows the relationships between elements in the UI. You can probably now appreciate how most, if not all, UI decisions inherently stem from UX considerations.
It is implored that you can now appreciate how expensive a UX research process can be. But don't sweat too much—generally, it is better to launch with decent coverage and constructive outlook, than supposedly exhaustive coverage but obstinate outlook.
Describe three common workflows within your application. Explain why those workflows were chosen over alternatives with regards to improving the user's overall experience in the context of an AI application.
User Interface
Beautiful UI design is the bare minimum now. With AI-assisted design and coding tools at your disposal, there's no excuse for a mediocre-looking application. Default component library styling (e.g., unstyled shadcn/ui defaults) isn't acceptable: you must customize and thoughtfully design your interface. If you're not confident in your design abilities, use AI to help you design. Some useful resources:
- Impeccable Style (opens in a new tab) (a guide to writing better CSS and improving your UI aesthetics)
- Design Prompts (opens in a new tab) (prompts and patterns for using AI to improve your UI/UX design).
That said, you still own your design, however you got there: handing it entirely to AI tends to produce something that looks and feels sloppy. As with poor AI-generated contents, sloppy design evokes a sloppy image and trust of your brand in your consumers.
UI is the space where interactions between humans and computers occur to achieve their objectives. You'd probably already be familiar with some of these concepts from Assignment 1, but UI design is more than just typography or color theory. Incorporating AI raises a few considerations specific to this assignment.
Ethics is a big deal here: be transparent about what your novel technology can and cannot do, cite third-party sources where possible, and hedge or admit uncertainty rather than overclaiming; this matters even more if your app touches sensitive fields like health, where an overconfident recommendation is as good as a misjudgement. For every recommendation, make sure the user ultimately has control over their decisions; don't assume, automate, or commit critical actions on the user's behalf, and instead give them the ability to accept, reject, or give feedback on your recommendation. For every intelligent UI element or feature, remember: the intelligence is artificial, and humans should always get the last say.
The same goes for privacy and data collection. Be transparent about what, when, and where you are collecting analytics or heuristics from a user's interactions, and give users the choice to opt out of certain intelligent features (think app permission settings on iOS and Android). None of this is unique to AI apps (these considerations apply to any app), but they matter more here, since AI features often touch sensitive data and take sensitive actions.
The basic principles of UI design still apply in this milestone: the above is meant to flag what's specific to building AI apps.
Show and explain considerations/decisions in your UI that were specially made for an app that leverages AI. Provide examples, citations, or justifications where necessary. You may also show different prototypes and outline their trade-offs.
In this milestone, we want you to consider how your interface is designed now that you have AI powering it. Just like developing for mixed reality, wearables, or any other new technologies, surely there must be UI decisions you make that's unique to AI. For example, adding disclaimers for AI-generated content, using suggest-accept for AI suggestions instead of automatic edits, or implementing suggestions, recommendations, or auto-fills for forms and fields. Google has a great guidebook to designing AI apps that centers around humans (opens in a new tab). We want to see how AI shapes and affords your app's interface. Oh, and of course, simply adding disclaimers or privacy notices won't score anything for you in this milestone because it's too trivial. You should use this milestone to show off the cool ones, and discard the trivial ones. If you have any questions about this milestone, feel free to contact the teaching staff to clarify or post to the Coursemology forum.
Phase 5: Launch
Landing Page
As with Assignment 1, waiting until your app is "ready" to launch a product site is usually a mistake. In this milestone, you're building a real product site that showcases your product: one that explains why it exists, informs about its features, and includes a clear call-to-action. Some decent examples to draw inspiration from include Remix (opens in a new tab), which pairs a clean, focused landing page with clear developer messaging.
Since this site ties into your later launch milestone, it's worth thinking about how it appears when shared: the thumbnail, title, and description snippet people see on Telegram, WhatsApp, or LinkedIn before they even click, defined by the Open Graph Protocol. (opens in a new tab)
To optimize your site for discovery and sharing, pay attention to:
- Search Engine Optimization (SEO): proper semantic HTML headers, page titles, and meta descriptions so your site is easily indexable.
- Social Sharing Previews: implement Open Graph Protocol (opens in a new tab) tags for attractive preview cards on platforms like Telegram, WhatsApp, LinkedIn, and X; consider Vercel's Dynamic OG Image (opens in a new tab) Generation for extra polish.
- Social Integrations: not strictly required, but social logins (e.g., Google, Facebook, GitHub) or share widgets like Facebook Social Plugins (opens in a new tab) and X for Websites (opens in a new tab) can lower friction for new users.
Create a landing page for marketing purposes with the following sections: hero, features, pricing section. Your landing page must be optimized for search engines (SEO) and support attractive social sharing previews (using the Open Graph Protocol (opens in a new tab)).
Analytics
You should understand how your application is actually being used. Many use Google Analytics 4 (opens in a new tab), Hotjar (opens in a new tab), Microsoft Clarity (opens in a new tab), and other analytics platforms to collect events, pageviews, and track specific user actions like button presses or form interactions.
With AI-assisted coding, integrating analytics is now very straightforward. The focus of this milestone is therefore less on the integration itself and more on what you can learn from the data: you should be able to articulate what your analytics tell you about user behavior, and how these insights inform your product decisions. Before relying on the report, verify events in a realtime or debug view and account for client-side route changes in single-page apps.
Embed an analytics tool in your application and give us a screenshot of the report. Explain what insights you derived from the data and how they influenced (or would influence) your product decisions. Analytics platforms usually report updates after some time (daily, weekly, etc.), so ensure that you embed and run the tracker early to receive enough and meaningful tracking information before the submission deadline.
Launch Campaign
Once your product is ready, it's time to announce it. You'll need a launch time, a plan for what happens leading up to it, and (most importantly) a plan for retaining users after the launch.
Round up your closest connections, forums, social media (Facebook, Instagram, etc.), communities (Reddit, X, HackerNews, Discord, etc.), and any influencers who might help spread the word. Decide what content best represents your product, draft it (text, video, banners, depending on the medium), mock it up, and revise until your team agrees on the best launch announcement, so that everything goes out together at launch.
Some products choose a soft launch (opens in a new tab) instead: think Clubhouse (opens in a new tab) or Arc Browser's (opens in a new tab) invite-only approach. This is useful for beta-testing and gathering feedback with a smaller, more manageable set of users, and creates a sense of exclusivity, but it can also slow acquisition; Arc's CEO has talked about losing up to 80% of sign-ups to their waitlist (opens in a new tab). Weigh the trade-off carefully; it will shape your product's acquisition curve.
After launch, keep monitoring your analytics, have technicians ready to fix bugs or restart servers if needed, and have customer support ready to answer questions. User acquisition takes repeated effort: don't coast once your active user curve ticks upward; start planning your next release or announcement before it plateaus.
Product Hunt, backed by Y Combinator (opens in a new tab), is a well-known place to launch new products (past highlights include Notion (opens in a new tab), Obsidian (opens in a new tab), Otter (opens in a new tab), and BeReal (opens in a new tab)), and has notable AI product launches (opens in a new tab) worth studying. To launch there, you'll need a product site, logo, first comment, demo video, and marketing pitch: see their guide on what to include (opens in a new tab) and their guide on preparing for a launch (opens in a new tab).
Here's a sample launch checklist Yangshun and his team used to launch Docusaurus 2.0 in 2022 (opens in a new tab) which contributed to Docusaurus 2.0 attaining the Product of the Day (opens in a new tab) award.
Even though this milestone frames things around a Product Hunt launch, the skills involved will be useful for your final project too.
Assume you were launching on Product Hunt. Come up with content and marketing materials that you will use for your Product Hunt submission. You may even want to launch on Product Hunt for real if you think your product is ready.
Phase 6: Going Above and Beyond (Optional)
Completing milestone(s) described in this section may contribute to the 30% coolness factor.
If these suggestions don't suit your application, you can still score full coolness points with ideas of your own. Blindly adding these technologies just to check a box won't earn you any credit. It's about using them in creative ways that make your application more desirable to use.
Some techniques once considered "above and beyond," such as basic embeddings/RAG and SEO/social sharing, are now increasingly standard practice (see Phase 3 and Phase 5 for more context). The suggestions below represent genuinely advanced techniques.
Advanced RAG
While basic RAG is increasingly common in modern AI applications (see Phase 3), there are more sophisticated retrieval techniques that can significantly improve the quality of your AI system's outputs:
- Graph RAG: Instead of treating your knowledge base as flat chunks of text, Graph RAG builds a knowledge graph that captures relationships between entities, enabling more nuanced and contextually-aware retrieval.
- Hybrid search: Combining dense vector search with sparse keyword search (e.g., BM25) for more robust retrieval across different query types.
- Agentic RAG: Using agents to dynamically decide what to retrieve, when to retrieve, and how to synthesize information from multiple sources.
Implement an advanced RAG technique (e.g., Graph RAG, hybrid search, agentic RAG) in your application. Explain the technique, why it was chosen, and demonstrate its impact compared to basic RAG.
Multi-Agent Systems and Agent Orchestration
For complex tasks, a single agent may not suffice. Multi-agent systems involve multiple specialized agents working together, each responsible for different aspects of a task. Agent orchestration involves coordinating these agents: managing their communication, task delegation, and result aggregation. Here’s a good page that talks more about such patterns (opens in a new tab).
Implement a multi-agent system or agent orchestration pattern in your application. Describe the architecture, the role of each agent, and how they coordinate to accomplish tasks.
Model Context Protocol (MCP)
Model Context Protocol (MCP) (opens in a new tab) is an open protocol that standardizes how AI applications connect to external data sources and tools. By implementing MCP, your application can expose its capabilities as tools that other AI systems can discover and use, or consume tools provided by other MCP-compatible services.
This is particularly valuable for building interoperable AI systems that can work together across different platforms and services.
Implement MCP in your application: either as an MCP server (exposing your app's capabilities) or an MCP client (consuming external MCP services). Explain how MCP integration enhances your product's value. Please integrate MCP only if you have a compelling reason to.
Assessment Scheme
The grading of the assignment is divided into two components: satisfying the compulsory milestones (70%) and the coolness factor (30%). Excluding Milestone 0, there are 20 compulsory milestones in total. Milestones 1, 2, 3 and 14 are worth 2.5% each. The rest are worth approximately 3.75% each.
The remaining 30% will be awarded based on the relative outcomes for the various teams. The top team might be awarded up to 30%, while the worst performing team less than 5%. The optional milestones can also contribute to this.
Overall, the AI application assignment is worth 20% of your final grade.
Mode of Submission
The following will need to be both submitted to Coursemology (under "Assignment 3: Artificial Intelligence Application") and included in your GitHub repository:
- A write-up,
group-<number>-milestones.pdf, containing your answers to all compulsory milestones that require written answers. Categorise your answers by the milestones they belong to. Please make sure that the URL for your live application and link to the public GitHub repository is clearly stated in the write-up for the convenience of the teaching staff. - A one/two-page pitch of your application,
group-<number>-pitch.pdf, pitching your application to the teaching staff, i.e., convince us that your application is so good that it deserves all 30% of the coolness points. Restriction: no longer than 2 A4 sides.
The following will only need to be included in your GitHub repository:
-
A
README.mdfile in the root directory. GitHub will automatically render it on your repository's front page. You may wish to style it using any of the supported markup languages. The file should contain:- The list of group members, including matriculation numbers, names and a description of the contributions of each member to the assignment.
- The name of your application.
- The URL to your application, i.e., your application must be accessible online somewhere.
- Set-up instructions for local testing (good to have).
- Resources you have used significantly to build your app (e.g., tutorials, source code references, design references, UI templates, etc.).
-
The code for your application. If your group is using Git submodules, make sure these submodules are also accessible by the teaching team (and that they are updated to reference the latest commits). You're encouraged to use monorepos (opens in a new tab) instead (e.g., via Turborepo (opens in a new tab)) so that you only have to submit one repository.
Not following the submission instructions (e.g., incorrect file naming) will result in the deduction of marks.
Feel free to reach out to the teaching staff if you need help at cs3216-staff@googlegroups.com or post to the Coursemology forum.
Good luck and have fun!