AI
Microsoft Agent Framework 1.0: Semantic Kernel Migration Guide
October 202613 min read

Not yet. Microsoft committed to keep fixing critical bugs and security issues in Semantic Kernel v1.x for at least one year after Agent Framework became generally available on 2 April 2026. Most new features, including the agent harness and CodeAct, are being built for Agent Framework instead.
Nothing replaces it directly; the Kernel simply goes away. An agent is created from a chat client with AsAIAgent in C# or Agent and as_agent in Python, and tools are passed in at creation. Message and content types come from Microsoft.Extensions.AI.
In Python, yes: semantic-kernel 1.38 and later adds an as_agent_framework_tool method that converts a KernelFunction, including prompt-based and vector search functions, into an Agent Framework tool. In C#, the usual port is to pass the same method to AIFunctionFactory.Create and drop the KernelFunction attribute. Either way, tools can move one at a time.
AsHarnessAgent in .NET and create_harness_agent in Python compose function invocation, per-call history persistence, todo tracking, plan and execute modes, file memory, tool approval middleware and OpenTelemetry by default. Context compaction is included but only becomes active when you supply token limits or a custom strategy. Each capability has its own opt-out flag.
It is pre-release. The Python package agent-framework-hyperlight was announced as alpha and is labelled beta in the docs, and it needs Linux with KVM or Windows with WHP. The .NET package is in preview and cannot restore until its sandbox dependency reaches nuget.org, so treat Microsoft's reported 52% latency and 64% token savings as a reason to benchmark your own workload.

Key Takeaway
A Microsoft Agent Framework migration from Semantic Kernel removes the Kernel: agents come from chatClient.AsAIAgent, tools from AIFunctionFactory.Create or plain Python functions, threads become agent-created sessions, and InvokeAsync becomes RunAsync. Semantic Kernel keeps critical and security fixes for at least a year after the April 2026 GA, so port tools first, then agents.
Consider a distributor in Jakarta whose ERP add-on answers stock questions through a Semantic Kernel ChatCompletionAgent. It works. Then Microsoft ships Agent Framework 1.0 on 2 April 2026, calls it the convergence of AutoGen and Semantic Kernel into a single supported platform, and every new sample, the agent harness and the CodeAct sandbox land there, not in Semantic Kernel. The question for that team is not whether to migrate, but how much of the code survives and in what order to move it.
This post answers both from the official Semantic Kernel and AutoGen migration guides, the harness and tool-approval documentation, and the BUILD 2026 announcement. It maps every Semantic Kernel concept to its Agent Framework replacement, shows the C# and Python code before and after, and covers the three things BUILD added that a hand-rolled Semantic Kernel agent never had: context compaction, standing tool approvals and default OpenTelemetry tracing.
Agent Framework is built by the same teams as Semantic Kernel and AutoGen and is positioned as the direct successor to both. It takes the simple agent abstractions from AutoGen and the enterprise parts of Semantic Kernel, meaning sessions, type safety, middleware and telemetry, and adds graph-based workflows for explicit multi-agent orchestration. It ships for .NET and Python, with Go in public preview. The 1.0 release of 2 April 2026 is the line Microsoft calls production-ready with stable APIs.
Semantic Kernel is not switched off. Microsoft committed in October 2025 to keep supporting Semantic Kernel v1.x, fixing critical bugs and security issues, for at least one year after Agent Framework became generally available, which puts the earliest end of that window in April 2027. The same post is just as clear that the majority of new features will be built for Agent Framework. So a working Semantic Kernel agent is not an emergency, but every month it stays there is a month without the harness, the new approval model and CodeAct. That makes a planned, tool-by-tool migration the sensible choice over either a rewrite or a freeze.
The single biggest change is that the Kernel disappears. In Semantic Kernel every agent depends on a Kernel instance, and the Kernel is where plugins, services and settings live. In Agent Framework an agent wraps a chat client directly, tools are passed in at creation, and message types come from Microsoft.Extensions.AI. Almost every other row in this table is a consequence of that one decision.
| Concern | Semantic Kernel | Agent Framework | What trips people up |
|---|---|---|---|
| Packages | Microsoft.SemanticKernel and Microsoft.SemanticKernel.Agents; pip semantic-kernel | Microsoft.Agents.AI plus Microsoft.Extensions.AI types; pip agent-framework, imported as agent_framework | The Python meta package pulls in several providers. Once you know yours, install agent-framework-core and only the provider packages you use. |
| Creating an agent | A Kernel plus ChatCompletionAgent, OpenAIAssistantAgent or AzureAIAgent | chatClient.AsAIAgent returning AIAgent; one ChatClientAgent over any IChatClient; in Python Agent(client=...) or client.as_agent() | Three service-specific agent classes collapse into one. Code that switches on the agent type has nothing left to switch on. |
| Conversation state | The caller picks an AgentThread subclass: ChatHistoryAgentThread, OpenAIResponseAgentThread, AzureAIAgentThread | agent.CreateSessionAsync or agent.create_session returns an AgentSession | AgentSession has no delete method. If you must delete hosted history, keep the provider conversation IDs yourself. |
| Tools | KernelFunction attribute, a KernelPlugin, and kernel.Plugins.Add | AIFunctionFactory.Create on a plain method, or a plain Python function, passed in tools | There is no plugin concept. Group related functions in a class if you like and pass the bound methods. |
| Running | InvokeAsync returns an async stream of AgentResponseItem | RunAsync or run returns one AgentResponse with Text and Messages | Loops that walk the items to find the final answer should read response.Text; tool calls and results are in Messages. |
| Streaming | InvokeStreamingAsync, invoke_stream | RunStreamingAsync returning AgentResponseUpdate; in Python run with stream=True | Python has no separate streaming method any more, only the flag. |
| Model options | OpenAIPromptExecutionSettings wrapped in KernelArguments | ChatClientAgentRunOptions over ChatOptions; in Python a TypedDict in options or default_options | MaxTokens becomes MaxOutputTokens, and ChatClientAgentRunOptions only applies to ChatClientAgent. |
| Dependency injection | services.AddKernel, then a Kernel injected into every agent | Register AIAgent itself, for example as a keyed singleton | The Kernel registration can go once no Semantic Kernel code still resolves it. |
Read the table as a dependency order. Tools have no dependency on the agent type, so they move first. Agent creation and sessions move together, because a session belongs to the agent that created it. Options and dependency injection follow last and are mostly search and replace.
Here is a stock-lookup agent written the Semantic Kernel way, kept close to the patterns in the official guide. Count the objects that exist only to carry one function and one conversation: the KernelFunction, the KernelPlugin, the Kernel, the settings wrapped in KernelArguments, and a thread class the caller has to match to the provider.
// Before: Semantic Kernel
using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.Agents;
KernelFunction onHand = KernelFunctionFactory.CreateFromMethod(InventoryPlugin.GetOnHand);
KernelPlugin plugin = KernelPluginFactory.CreateFromFunctions("Inventory", [onHand]);
Kernel kernel = BuildKernel(); // every SK agent needs one
kernel.Plugins.Add(plugin);
OpenAIPromptExecutionSettings settings = new() { MaxTokens = 1000 };
AgentInvokeOptions options = new() { KernelArguments = new(settings) };
ChatCompletionAgent agent = new() { Instructions = StockInstructions, Kernel = kernel };
// The caller has to know which thread class matches the provider.
AgentThread thread = new OpenAIResponseAgentThread(responsesClient);
await foreach (AgentResponseItem<ChatMessageContent> item
in agent.InvokeAsync(question, thread, options))
{
Console.WriteLine(item.Message);
}
public class InventoryPlugin
{
[KernelFunction] // required, or the model never sees it
public static int GetOnHand(string itemCode, string warehouseCode) => /* ERP query */ 0;
}The Agent Framework version keeps the same chat client and the same method. AsAIAgent registers the tool in the same call that creates the agent, the Description attribute replaces KernelFunction and is optional, and the agent decides which session type its provider needs.
// After: Microsoft Agent Framework
using System.ComponentModel;
using Microsoft.Agents.AI;
using Microsoft.Extensions.AI;
// chatClient is any IChatClient: Azure OpenAI, OpenAI, Foundry, Ollama...
AIAgent agent = chatClient.AsAIAgent(
instructions: StockInstructions,
tools: [AIFunctionFactory.Create(GetOnHand)]);
// The agent creates the right session type for its provider.
AgentSession session = await agent.CreateSessionAsync();
// MaxTokens is now MaxOutputTokens, on ChatOptions.
ChatClientAgentRunOptions options = new(new() { MaxOutputTokens = 1000 });
// One AgentResponse instead of an async stream of items.
AgentResponse response = await agent.RunAsync(question, session, options);
Console.WriteLine(response.Text); // tool calls and results are in response.Messages
[Description("On-hand quantity for an item code in one warehouse.")] // optional
static int GetOnHand(string itemCode, string warehouseCode) => /* ERP query */ 0;The behavioural difference worth testing is the return shape. InvokeAsync yielded several items and most Semantic Kernel code printed each one; RunAsync returns one AgentResponse whose Messages list holds the tool calls, function results, reasoning updates and final answer together. Any code that logged each item for an audit trail now needs to walk response.Messages instead, or it will silently log only the final text.
Do not port hosted-thread cleanup line for line. Semantic Kernel threads had a delete method; AgentSession does not, because not every provider supports hosted history or deleting it. If your retention policy requires deleting conversations, record the session's conversation ID when you create it and delete through the provider's own SDK.
The Python side changes the same things, with one useful escape hatch. Since semantic-kernel 1.38, a KernelFunction has an as_agent_framework_tool method that turns it into an Agent Framework tool, including functions built from prompt templates and vector-store search functions. That lets a team move the agent first and port tools one pull request at a time.
# pip install agent-framework (imported as agent_framework)
from typing import Annotated
from agent_framework import tool
from agent_framework.openai import OpenAIChatClient
from semantic_kernel.functions import kernel_function # still installed during the move
# Ported: a plain function. Name -> tool name, docstring -> description.
@tool(approval_mode="never_require") # read-only lookup, safe to auto-run
def get_on_hand(
item_code: Annotated[str, "ERP item code, e.g. BRG-00412"],
warehouse_code: Annotated[str, "Warehouse code, e.g. SBY-01"],
) -> int:
"""On-hand quantity for an item code in one warehouse."""
...
# Not ported yet: the existing Semantic Kernel function, bridged.
# Needs semantic-kernel 1.38 or later.
@kernel_function(name="reorder_hint", description="Suggest a reorder quantity from min/max policy")
def reorder_hint(item_code: str) -> str:
...
reorder_tool = reorder_hint.as_agent_framework_tool()
agent = OpenAIChatClient().as_agent(
instructions="Answer stock questions for the Surabaya warehouse.",
tools=[get_on_hand, reorder_tool],
default_options={"max_tokens": 1000}, # was OpenAIPromptExecutionSettings + KernelArguments
)
session = agent.create_session() # was ChatHistoryAgentThread()
response = await agent.run(question, session=session, options={"max_tokens": 500})
print(response.text)
# Streaming is a flag on run(), not a separate invoke_stream() method.
async for update in agent.run(question, session=session, stream=True):
print(update.text, end="")Two details in that code are deliberate. The plain function becomes a tool whose name is the function name and whose description is the docstring, so a vague docstring is now a vague tool description the model reads. And options move into a typed dictionary: default_options at creation, options per call, while tools and instructions stay as ordinary keyword arguments.
AutoGen users face a bigger conceptual shift than Semantic Kernel users, because orchestration changes model, not just names. AutoGen pairs an event-driven core with a high-level Team; Agent Framework centres on a typed, graph-based Workflow that routes data along edges and activates executors when their inputs are ready. The mappings that matter most:
One limit to plan around: AutoGen offers embedded and experimental distributed runtimes, while the AutoGen guide says Agent Framework focuses on single-process composition today, with distributed execution planned. A system that depends on AutoGen's distributed runtime should not be first in the migration queue.
The harness is the part of Agent Framework that a Semantic Kernel team would otherwise have had to build itself. AsHarnessAgent in .NET, from the Microsoft.Agents.AI.Harness package, and create_harness_agent in Python return an ordinary agent with a pipeline already composed: function invocation with an iteration limit, history persisted after each model call, todo tracking, plan and execute modes, session file memory, tool approval middleware and OpenTelemetry instrumentation. Here it drives a cycle-count reconciliation where lookups run freely but stock adjustments wait for a supervisor.
using Microsoft.Agents.AI;
using Microsoft.Extensions.AI;
// Lookups run freely. Anything that writes to the ledger waits for a person.
AIFunction onHand = AIFunctionFactory.Create(GetOnHand);
AIFunction postAdjustment = new ApprovalRequiredAIFunction(
AIFunctionFactory.Create(PostStockAdjustment));
AIAgent agent = chatClient.AsHarnessAgent(new HarnessAgentOptions
{
Name = "stock-reconciler",
HarnessInstructions = "Use tools deliberately and report verified results.",
ChatOptions = new ChatOptions
{
Instructions = "Reconcile cycle-count sheets against ERP on-hand quantities.",
Tools = [onHand, postAdjustment],
},
// Compaction only switches on when the harness has a budget to work against.
MaxContextWindowTokens = 128_000,
MaxOutputTokens = 16_384,
// Web search is added by default where the client supports it.
// An internal ERP agent has no reason to browse.
DisableWebSearch = true,
});
AgentSession session = await agent.CreateSessionAsync();
AgentResponse response = await agent.RunAsync(countSheetSummary, session);
// A posting attempt comes back as a request, not as a result.
List<ToolApprovalRequestContent> pending = Pending(response);
while (pending.Count > 0)
{
List<AIContent> answers = [];
foreach (ToolApprovalRequestContent request in pending)
{
var call = (FunctionCallContent)request.ToolCall;
bool approved = await AskSupervisorAsync(call.Name, call.Arguments);
answers.Add(request.CreateResponse(approved));
}
// Same session: approval responses are bound to the requests recorded in it.
response = await agent.RunAsync(new ChatMessage(ChatRole.User, answers), session);
pending = Pending(response);
}
static List<ToolApprovalRequestContent> Pending(AgentResponse r) =>
r.Messages.SelectMany(m => m.Contents).OfType<ToolApprovalRequestContent>().ToList();Three behaviours in that code are easy to get wrong. Compaction, the feature that summarises chat history mid-loop so a long tool loop does not overflow the context window, is only active when you supply token limits or a custom strategy, so a harness without MaxContextWindowTokens does not compact. Web search is added by default wherever the chat client supports it, which is why the example turns it off. And the approval loop must reuse the same session, because approval responses are bound to the requests recorded in it; a fabricated or replayed approval that matches no pending request is ignored.
The approval layer has two parts that the docs keep separate. ApprovalRequiredAIFunction in .NET, or approval_mode set to always_require in Python, marks a tool as needing approval. The harness then adds ToolApprovalAgent middleware on top, which queues multiple requests, remembers standing approvals from earlier answers and can apply auto-approval rules you supply. Setting DisableToolAutoApproval removes only that middleware; it never removes the requirement from a tool you wrapped.
Every harness capability has its own opt-out, so trim rather than accept the bundle. In .NET the switches include DisableWebSearch, DisableFileMemory, DisableTodoProvider, DisableAgentModeProvider, DisableAgentSkillsProvider, DisableOpenTelemetry and DisableCompaction; Python has disable_web_search, disable_file_memory, disable_todo, disable_mode and disable_compaction. An ERP back-office agent rarely needs web search or skills, but it almost always wants tracing left on.
CodeAct changes the tool-calling loop itself. Instead of a model, tool, model, tool chain, the model writes one short Python program that calls the tools through call_tool, and the program runs in a fresh, isolated Hyperlight micro-VM. Microsoft's BUILD benchmark on a multi-step workload went from 27.81 seconds and 6,890 tokens with traditional tool calls to 13.23 seconds and 2,489 tokens with CodeAct, the reported 52.4% latency and 63.9% token reduction. Treat that as one vendor-run workload, not a promise; the docs suggest timing both modes on your own tasks.
# pip install agent-framework-hyperlight --pre
# Needs x86-64 Linux with KVM or AMD64 Windows with WHP, Python 3.10 to 3.14.
from agent_framework import Agent, tool
from agent_framework.hyperlight import HyperlightCodeActProvider
@tool
def get_stock_card(item_code: str, month: str) -> list[dict]:
"""Every stock movement for one item in one month (read-only)."""
...
@tool
def get_standard_cost(item_code: str) -> float:
"""Current standard cost of an item (read-only)."""
...
@tool(approval_mode="always_require")
def post_revaluation(item_code: str, new_cost: float) -> str:
"""Post a cost revaluation journal. Changes the books."""
...
codeact = HyperlightCodeActProvider(
# Hidden from the model as direct tools; reachable only via call_tool(...)
# inside one execute_code program. These still run in the HOST process.
tools=[get_stock_card, get_standard_cost],
approval_mode="never_require",
)
agent = Agent(
client=client,
name="CostAnalyst",
instructions=(
"To analyse many items, write one program that uses call_tool(...) "
"and end it with print(...). Never revalue without asking."
),
tools=[post_revaluation], # stays a first-class tool, approved per call
context_providers=[codeact],
)The split in that example is the documented rule of thumb. Cheap, deterministic, read-only tools go on the provider, where the model can chain dozens of calls in one execute_code turn; that is exactly the shape of a month-end stock-card analysis. The revaluation tool stays a direct tool with approval_mode set to always_require, because approval in CodeAct applies to the whole execute_code invocation, not to each call_tool inside it.
CodeAct is pre-release and narrower than it looks. The Python package agent-framework-hyperlight was announced as alpha and the docs now call it beta; it needs x86-64 Linux with KVM or AMD64 Windows with WHP, and Python 3.10 to 3.14. The .NET package is in preview and, per the docs, fails to restore until its Hyperlight.HyperlightSandbox.Api dependency is published to nuget.org. And the sandbox does not contain your tools: call_tool is a bridge back to the host process, so a tool registered on the provider runs with the host's credentials and network.
Many Indonesian companies already run on the Microsoft stack, and Azure's Indonesia Central region in Jakarta makes keeping the surrounding resources in-country straightforward. For such a team, an order that keeps the system shippable at every step looks like this:
The order matters because each step is independently testable. A tool ported in step two returns the same rows to the old Semantic Kernel agent and the new one, so a mismatch is caught before the agent changes, not after.
Moving from Semantic Kernel to Microsoft Agent Framework is mostly deletion: the Kernel, the plugin wrapper and the thread classes go, and the chat client, the methods and the model stay. The rule to carry is to move tools before agents, mark every side effect for approval before the harness touches it, and treat CodeAct's headline numbers as a reason to measure, not a result.
Sources and further reading