Back to notes

August 13, 2026 · 17 min read

From Packet Capture to Dart AOT: Unpacking A/B Routing in a Flutter iOS App

How I went from a suspicious difference in app behavior to the client-side A/B routing logic using USB packet capture, remote-config decoding, and Dart AOT analysis.

It started with something that did not quite add up

While testing a Flutter iOS app, I noticed that changing users could also change parts of the content and new-user benefits, even though the app version stayed the same. The app also fetched remote configuration during startup, which made the difference feel like more than ordinary personalization.

One possibility came to mind: perhaps the app showed a safer, more limited experience during review, then let the server route regular users to another experience after approval. That second side might expose sensitive material that reviewers never saw, potentially including sexual content or other material that would not comply with store policy.

At that point, however, this was only a suspicion. I first needed to answer smaller, concrete questions. Was there a stable routing decision in the client? What inputs did it use? How did the result reach the server?

This kind of investigation is easy to derail with visible but weak clues. Two users may have different balances. One request may go through a proxy. Device language and time zone may differ. Every one of those details looks suspicious in isolation, but without controlled comparisons they are leads, not answers.

I eventually split the problem into four layers:

Can I observe the traffic?
  → What does the app actually send?
  → How does the client compute its routing value?
  → Does the server maintain additional user state?

Everything described here was tested on devices and samples under my control. I have removed live credentials, real device identities, product-specific names, and details that would reproduce production requests. This article is about understanding the mechanism, not bypassing it.

Why I kept digging

Remote config, feature flags, and staged rollouts are normal engineering tools. We use them to release features gradually, adapt behavior by region, or disable a broken module without waiting for a new binary. Dynamic delivery by itself is not suspicious.

The interesting question is whether it creates an undisclosed difference between the core experience shown to reviewers and the one shown to regular users.

The possible flow I had in mind looked like this:

One IPA may show different experiences in review and regular-user contexts

Apple’s App Review Guidelines are clear on this point. Apps should not hide dormant or undocumented functionality from review, major features and product changes must be accessible to reviewers, and overtly sexual or pornographic material is not accepted. The guidelines also warn that attempts to deceive the review process can lead to removal from the store and termination from the Developer Program.

Still, an app that can switch screens remotely is not automatically evading review. Most apps with remote config use it for legitimate reasons. To support the stronger claim, I would need evidence that the split was tied to a review context and that the other side actually exposed undisclosed, noncompliant content.

First, define what would count as evidence

To avoid talking myself into a conclusion after finding a few odd field names, I wrote down what each kind of observation could and could not establish:

What I observeWhat it supportsWhat it does not prove by itself
Remote switches and two UI paths existThe app can change experiences dynamicallyThe switches are used to evade review
Different identities receive different config or contentThe server performs user routingApp Store review is the routing criterion
Review-related contexts consistently see the safe side while regular contexts see sensitive contentStrong support for the review-evasion hypothesisWhether the material is illegal still depends on jurisdiction and the actual samples

To make a convincing case, I would want to connect at least these points:

  1. The same binary can produce two observable content experiences.
  2. The client contains a clear routing variable and call chain.
  3. Remote config or the user API can change that variable.
  4. The result remains reproducible after controlling for language, time zone, network path, and old versus new identities.
  5. Sensitive content comes from the alternate route rather than user uploads, stale cache, or a one-off recommendation.
  6. A request-and-response timeline shows that the change really occurs around the review period.

That framing also shaped the investigation. I did not begin by patching the condition to force the other side. Instead, I tried to record how the unmodified app made the decision and how the server responded. Original behavior is much stronger evidence than a result I manufactured myself.

Getting Flutter traffic into the proxy

I did not start with disassembly. The first job was making the network traffic observable.

Most native iOS apps use the system networking stack, so setting an HTTP proxy is often enough for tools such as Proxyman or Charles. A Flutter app may establish connections through Dart’s own networking implementation and ignore the system proxy. That was exactly the symptom here: other apps appeared in the proxy, while the target app looked as if it never connected at all.

The test device was connected to the Mac by cable, so I built this capture path:

Flutter iOS traffic routed through Frida and USB forwarding into Proxyman

The important part was not the tool name. It was verifying every hop:

  1. Did the app actually resolve the target hostname?
  2. Which IP address and port did the socket use?
  3. Was the forwarded USB port reachable?
  4. After handling TLS trust, did the full request reach the proxy?
  5. Did changing the network path leave the business response unchanged?

Once the chain worked, I could reliably inspect the target HTTPS traffic. Headers, bodies, request order, and responses no longer had to be guessed.

The custom header looked encrypted, but was not

Business requests included a custom header. After reconstruction, its contents were simply JSON. To keep the sample anonymous, the field names and values below are semantic aliases rather than the product’s original names:

{
  "locale": "en",
  "channel": "Default",
  "packageName": "SampleAppIOS",
  "appId": "<device-scoped-id>",
  "token": "<user-token>",
  "deviceType": "2",
  "version": "x.y.z",
  "timeOffset": "0",
  "useVpn": ""
}

The transformation was not cryptographically meaningful:

Compact JSON
  → Base64
  → insert two random strings at fixed positions

Decoding simply removed those two strings and Base64-decoded the remainder. This can keep plaintext out of casual logs and interfere with basic string matching, but it provides no real confidentiality.

That was a useful reminder: when a value looks unreadable, do not jump straight to AES or RSA. Compare length, alphabet, padding, and stable regions across samples first. It is often much quicker to distinguish encryption from encoding or light obfuscation that way.

The remote config contained another JSON layer

During startup, the app fetched remote configuration. The outer response looked roughly like this:

{
  "message": "Operation success.",
  "success": true,
  "encryptStatus": 0,
  "data": {
    "content": "<obfuscated-base64>"
  }
}

The encryptStatus value of 0 was easy to misread. content was still not directly readable. It used the same style of inserted noise: remove two random segments, Base64-decode the result, and parse the first JSON object. Even then, customJson remained a JSON string and needed a second parse.

A simplified Dart decoder looks like this:

Map<String, dynamic> decodeConfigContent(String content) {
  final cleanBase64 =
      content.substring(0, 10) +
      content.substring(20, content.length - 20) +
      content.substring(content.length - 10);

  final decodedText = utf8.decode(base64Decode(cleanBase64));
  final result = jsonDecode(decodedText) as Map<String, dynamic>;

  result['customJson'] = jsonDecode(result['customJson'] as String);
  return result;
}

The key values in the response at the time were:

{
  "featureGateA": "0",
  "controller": "sample_1",
  "customJson": {
    "variantOpenMode": 1,
    "variantAppName": "sample",
    "variantChannel": "Default",
    "selfiePrice": 10,
    "adWaitSeconds": 180,
    "safetyThreshold": 0.5,
    "rewardAdLimit": 3
  }
}

The full object also contained a file domain, server time, resolution presets, video pricing, crop dimensions, update links, and currently empty H5 and push settings. Dynamic tokens, keys, and timestamps changed between requests, so I excluded them here and did not treat them as stable configuration.

Ruling out the most tempting explanations

Once traffic was visible, I ran controlled comparisons before touching the client.

The variables included:

  • old versus newly generated appId values;
  • locale and time zone;
  • direct traffic versus different network exits;
  • User-Agent and ordinary HTTP headers;
  • remote config returned to a review-like user and a regular user.

New identities received the same new-user benefits across several combinations of locale, time zone, and network path. The existing identity continued to preserve its historical state. After removing dynamic tokens, keys, and timestamps, the remote config returned to both identities was identical.

That supported four limited conclusions:

  1. The remote config endpoint was not returning different config to those two identities at that time.
  2. Locale, time zone, and exit IP did not independently explain the observed difference.
  3. Balance or welcome-credit differences looked more like account history than a review flag.
  4. An existing appId may be associated with server-side historical state.

This does not mean locale, IP, and device identity are never used. It only means that changing those inputs individually did not change the result in this test. The server could still combine them with account history, device records, or other risk signals.

Then it was time for Dart AOT

Packet capture could show what the client sent, but not how those parameters were computed. For that, I had to move into the Flutter AOT output.

The sample used a relatively recent Dart version. Most business logic lived in the Dart snapshot inside App.framework, not in a convenient set of Objective-C or Swift symbols. Searching Mach-O strings exposed endpoint paths and field names, but not where or in what order they were used.

I parsed the Dart snapshot with Blutter and made a temporary adaptation for the snapshot regions in an iOS Mach-O. Full analysis did not complete because this particular build used uncompressed pointers, but --no-analysis mode still recovered many class names, function names, object-pool strings, and relative addresses. Combining those results with ARM64 disassembly gradually made the call graph readable.

The useful output was not a piece of pseudo-source that merely looked convincing. It was three kinds of evidence that could check one another:

InformationWhat it tells me
Dart string poolWhich endpoint and field names really exist in the binary
Recovered class and function namesWhether a code region belongs to config, user, or networking logic
ARM64 calls and branchesIn what order values are read and how they affect control flow

Following those three sources led to a boolean getter on the user object. I call it variantLevel here to keep the original identifier anonymous.

The real client-side routing entry was variantLevel

The getter’s original name looked as though it represented a numeric level, but it actually returned true or false. It read these inputs:

localChannel       locally stored channel
variantChannel     channel required by remote config
userChannel        current user's channel, defaulting to Default
rechargeStatus     whether the user has recharged
variantOpenMode    routing mode
hiddenGate         hidden switch controlled by config or server

Reduced to readable Dart, the assembly was approximately:

bool calculateVariantLevel({
  required String savedChannel,
  required String typeChannel,
  required String musicType,
  required int rechargeStatus,
  required int openStatus,
  required bool hiddenGate,
}) {
  final channelCondition =
      !containsIgnoreCase(typeChannel, musicType) ||
      !containsIgnoreCase(typeChannel, savedChannel) ||
      rechargeStatus == 1;

  switch (openStatus) {
    case 0:
      return channelCondition;
    case 1:
      return channelCondition && hiddenGate;
    case 2:
      return hiddenGate;
    default:
      return false;
  }
}

containsIgnoreCase is a case-insensitive containment check, not strict string equality. That small distinction can change the result for some channel values.

Config parsing also computed a hidden switch:

final hiddenGate =
    featureGateA != '1' &&
    appName == controller;

The routing parameter placed in the user-initialization request was the inverse of variantLevel. I call it variantType here:

final variantType = variantLevel ? 0 : 1;

Substituting the values from this sample:

openStatus = 1
featureGateA = 0
appName = sample
controller = sample_1

Because appName and controller did not match, the first config parse produced hiddenGate=false. With openStatus=1, both the channel condition and hidden gate had to pass, so variantLevel=false and the outgoing variantType became 1.

The server could make the client calculate again

The response handler for user initialization contained another important branch. If the server returned a synchronization flag, the client set the hidden gate to true and repeated the user request. I call that flag syncFlag.

Putting config evaluation and server feedback together makes the flow easier to see:

Remote config, user state, and a server sync flag determine the client routing value

This is why a single captured request can be misleading. The server does not merely receive variantType; it can use its response to change a client-side switch and trigger the calculation again.

Several suspicious-looking fields were not involved

The user response contained many fields that looked like identity labels: account type, chat state, profile state, balance, and subscription state. Looking only at the JSON, it would be easy to pick one and call it the review-user flag.

Static references and the user-binding routine told a different story. The values the client explicitly extracted and stored were mainly:

authentication token
nickname
avatar
balance

The AOT string pool also contained no direct blacklist, isReview, or reviewStatus field. User type, chat state, and profile state did not feed into variantLevel either.

The “login IP country” field was another tempting clue. It did exist in the app, but its references belonged to web-payment and regional selection code, not home-screen routing or user initialization. On the client side, it therefore could not support a claim that country-by-IP controlled the A/B side. Whether the server independently used the source IP remained invisible from the client binary.

What the client did and did not establish

At this point, the client-side behavior was fairly clear.

I could confirm from the client that:

  • remote config was parsed before login or user binding;
  • variantLevel was the central local routing boolean;
  • channel values, recharge state, routing mode, and the hidden gate contributed to it;
  • variantType was the inverse of variantLevel;
  • syncFlag=1 could change the hidden gate and trigger a retry;
  • common user-type fields in the response did not directly drive this decision.

But the client could not answer:

  • how the server database recorded a particular existing appId;
  • whether IP, locale, time zone, device history, and risk signals were combined server-side;
  • whether the server returned different content for the same variantType depending on identity;
  • when or by which backend job an account’s historical state was written.

So instead of saying that the client blacklisted a user on a particular line, the more accurate conclusion was:

The client generates a routing hint, variantType, from variantLevel. The server can still preserve state by identity and ask the client to retry with a changed gate. An old identity remaining on one side is more likely to reflect server-side state than a direct blacklist field read by the client.

Could this mechanism switch an App Store review side?

Structurally, most of the pieces required for pre- and post-review switching were present:

Required pieceStatusWhat it enables
Remote config fetched at startupConfirmedThe server can change runtime parameters
A client-side routing functionConfirmedChannel, account state, and hidden gate collapse into one result
A routing value in the requestConfirmedThe server learns the client’s current side
A server synchronization flagConfirmedThe server can ask the client to change its gate and retry
Historical state preserved by identityConsistent with experimentsThe same device identity can remain on one side
Automatic identification of App Store reviewersNot confirmedRequires review-period samples and server-side evidence
Stable delivery of noncompliant content to regular usersNot yet established end to endRequires preserved content responses and a repeatable timeline

This was not a simple if hard-coded into a page. Even if the App Store reviewed exactly the same IPA that users later installed, the backend could change what content the client requested through remote config, account history, and the synchronization flag. No new binary would be required.

That is why this pattern deserves attention. A reviewer sees the app at one time, on one network, under one identity. The remote system can later change the experience while the binary remains untouched. Static analysis may reveal that two paths exist, but not necessarily the backend conditions that select one of them.

Even so, this investigation did not prove that the server recognized Apple reviewers. The two test identities received identical remote config, and changing locale, time zone, or one exit IP did not independently change the new-identity outcome. The stronger signal was that the backend appeared to preserve state for an existing identity. The client sample could not show whether that state originally came from a review IP range, device characteristics, registration time, a manual backend flag, or another risk model.

If I continue this investigation, the most useful next step is not guessing more parameters. It is preserving a review-period timeline: for the same identity and the same binary, record config, synchronization flag, routing value, and content list before, during, and after review. Only a repeatable timeline can move the claim from “the architecture permits it” to “the app actually did it.”

A few wrong turns along the way

Treating balance as a review signal

One user had no balance while a new user received a welcome credit. At first glance, that looked exactly like an A/B difference. But balance also depends on signup time, promotions, spending, and account history. Without comparing content lists, config, and request parameters together, it remained a weak clue.

Trusting locale, time zone, and IP too early

Those are common review and risk inputs, but common does not mean proven in this sample. New identities still had to be tested one variable at a time to see whether any input independently changed the result.

String search is useful for drawing the first map, but it does not prove a call relationship. Finding an IP-country field does not mean it participates in routing. Finding a routing parameter does not explain why it becomes 0 or 1. Function ownership, object-pool references, and branch instructions still had to agree.

Calling obfuscation encryption

The custom header and config content were Base64 with random characters inserted at fixed positions. Calling that “decryption” in casual conversation is convenient, but it can send the investigation toward cryptographic keys that do not exist.

The order I would use next time

If I encounter another Flutter app with remote config and A/B routing, I would follow roughly the same sequence:

  1. Make the traffic visible and determine whether the app ignores the system proxy.
  2. Record the full startup request order rather than isolating one endpoint.
  3. Separate the outer response, encoded content, and nested JSON strings.
  4. Compare each suspicious variable independently.
  5. Remove dynamic tokens, keys, and timestamps before comparing responses.
  6. Trace a routing field backward to the function that generates it.
  7. Cross-check the Dart string pool, recovered function names, and ARM64 branches.
  8. Separate confirmed behavior, reasonable inference, and server-side unknowns.
  9. Prefer observation over modifying identities or forcing a route.

The number of tools matters less than whether every conclusion can answer three questions: where did it come from, how can it be reproduced, and what other explanation could still fit?

Looking back

I began with a narrow question: was this app using dynamic delivery to distinguish review users, and where was that decision made? By the end, no single function address was enough to summarize the answer.

The client did contain A/B routing. The important pieces were not in page code, but in startup config parsing, variantLevel, and variantType. The server could preserve historical state by identity and use syncFlag to affect the next client request. Together, those pieces were sufficient to support a design in which review sees a safe side and regular users later see another side.

But “sufficient to do it” is not the same as “proven to have done it.” What I could confirm was the routing mechanism and the boundary of remote control. I could not yet confirm exactly how the server recognized an App Store review identity, or whether the alternate content changed consistently around the review timeline.

That was the most important lesson from the investigation. Seeing client code does not mean seeing the whole system. Capturing a server response does not reveal how its underlying state was created.

A conclusion becomes convincing not when one field looks suspicious, but when static code, live traffic, and controlled experiments all point to the same call chain. Everything still unknown should remain honestly labeled as unknown.